e860167a-2cfb-4415-a19c-86b476b2a1d6
· 2026-07-29T20:05:06.013Z
· 67m11s wall-clock
| Target | /Users/rupulsafaya/Documents/GitHub/canaryone-demo |
| Model(s) | accounts/fireworks/routers/kimi-k3-fast, kimi-k3, moonshotai/Kimi-K3, moonshotai/kimi-k3 |
| Lanes | 10 (accounts/fireworks/routers/kimi-k3-fast@direct:fireworks, kimi-k3@direct:moonshot-intl, moonshotai/Kimi-K3@direct:nebius, moonshotai/kimi-k3@vercel:baseten, moonshotai/kimi-k3@vercel:fireworks, moonshotai/kimi-k3@vercel:moonshotai, moonshotai/kimi-k3@vercel:nebius, moonshotai/kimi-k3@openrouter:moonshotai/mxfp4, moonshotai/kimi-k3@openrouter:fireworks/fast, moonshotai/kimi-k3@openrouter:nebius/fp4) |
| Sessions | 120 = 4 tasks × 10 lanes × 3 repeats |
| Pass rate | 86/120 (72%) · 15 failed |
| Total spend | $10.53 |
canaryone is a local CLI. Point it at your codebase, and it runs your own tests against multiple LLM providers to compare cost and quality on real work — the LLM your test would normally call gets swapped for whichever provider we're benchmarking. Nothing leaves your machine.
For this run, canaryone:
/Users/rupulsafaya/Documents/GitHub/canaryone-demo.accounts/fireworks/routers/kimi-k3-fast@direct:fireworkskimi-k3@direct:moonshot-intlmoonshotai/Kimi-K3@direct:nebiusmoonshotai/kimi-k3@vercel:basetenmoonshotai/kimi-k3@vercel:fireworksmoonshotai/kimi-k3@vercel:moonshotaimoonshotai/kimi-k3@vercel:nebiusmoonshotai/kimi-k3@openrouter:moonshotai/mxfp4moonshotai/kimi-k3@openrouter:fireworks/fastmoonshotai/kimi-k3@openrouter:nebius/fp4.c1/db.sqlite and rendered this report.Terms you'll see below:
(model, provider) pair — a comparison target.N times to smooth out variance. "Attempt 3/3" = the third repeat of that session.The report's primary comparison metric is weighted $/pass — the dollars you'd spend to get one grounded pass on your workload. It penalises passes where the test succeeded but the model didn't do real work (e.g. narrated a plausible answer instead of grounding in tool results).
weighted $/pass = $/pass ÷ (judge score / 100)Raw $/pass counts every passing exit-code as equal. The judge score (0-100) is a composite of four sub-scores of 25 each — two computed deterministically from the wire log (Action, Efficiency) and two scored by the judge LLM (Grounding, Verification). A low judge score with a passing test means "the test passed but the model didn't really do the work" — the ⚠ badge flags this on scores under 50.
Weighted $/pass collapses "cheap" and "actually good" into one number so there's a single winner per run, not a Pareto curve.
Four headline cards summarising the run. Best value is the lane with the lowest weighted $/pass among those with ≥1 pass — the primary metric. Cheapest raw ignores trajectory quality; if its traj is < 50 it carries a ⚠ narrated flag (interpret with §4.5 heuristic). Pass rate is total sessions passed across all lanes. Spread is most-expensive weighted / best-value weighted — how far apart the extremes are on identical work.
<router>:<provider>. Same weights served through different routers = different lanes.passed / attempted for this lane.total spend / passed count. Ignores trajectory quality.⚠ on scores < 50 (usually means "test passed but the model didn't really do the work"; see §4.5 of the SPEC for the caveat around workloads that don't exercise tool_calls).$/pass ÷ (judge score / 100). Penalises narrated passes. Default sort ascending.Click any column header to sort. Green row = winner (lowest weighted $/pass).
| Lane | Model | Router | Pass | $/pass | Judge | Weighted $/pass | p50 lat. | p95 lat. |
|---|---|---|---|---|---|---|---|---|
| vercel:nebius | moonshotai/kimi-k3 | VRC | 6/12 | $0.0472 | 69 | $0.0683 | 5447ms | 303725ms |
| direct:moonshot-intl | kimi-k3 | DIR | 9/12 | $0.0686 | 87 | $0.0789 | 8693ms | 26144ms |
| vercel:moonshotai | moonshotai/kimi-k3 | VRC | 9/12 | $0.0957 | 90 | $0.1064 | 8087ms | 29543ms |
| vercel:fireworks | moonshotai/kimi-k3 | VRC | 9/12 | $0.0921 | 86 | $0.1071 | 5874ms | 16732ms |
| vercel:baseten | moonshotai/kimi-k3 | VRC | 11/12 | $0.0967 | 84 | $0.1151 | 5354ms | 32722ms |
| direct:nebius | moonshotai/Kimi-K3 | DIR | 6/12 | $0.1052 | 86 | $0.1224 | 5126ms | 296788ms |
| direct:fireworks | accounts/fireworks/routers/kimi-k3-fast | DIR | 7/12 | $0.1315 | 79 | $0.1664 | 2948ms | 8037ms |
| openrouter:nebius/fp4 | moonshotai/kimi-k3 | OR | 8/12 | $0.1578 | 80 | $0.1973 | 4110ms | 20309ms |
| openrouter:moonshotai/mxfp4 | moonshotai/kimi-k3 | OR | 12/12 | $0.1852 | 88 | $0.2104 | 7791ms | 28660ms |
| openrouter:fireworks/fast | moonshotai/kimi-k3 | OR | 9/12 | $0.2040 | 84 | $0.2428 | 9788ms | 28454ms |
One row per lane, one column per task. Cell color = weighted $/pass gradient (green cheapest → red most expensive). Cell text = raw $/pass for that (lane, task) combination. Right-most column shows the lane's overall weighted $/pass with judge score badge. Toggle to raw to color by raw $/pass instead — surfaces "cheap by naive metric" lanes that hide a low judge score.
| Lane | t01 | t02 | t03 | t04 | Weighted $/pass |
|---|---|---|---|---|---|
| vercel:nebius | $0.0503 | $0.0440 | — | — | $0.0683 (judge 69) |
| direct:moonshot-intl | $0.0471 | $0.0371 | $0.1217 | — | $0.0789 (judge 87) |
| vercel:moonshotai | $0.0756 | $0.0602 | $0.1514 | — | $0.1064 (judge 90) |
| vercel:fireworks | $0.0547 | $0.0536 | $0.1681 | — | $0.1071 (judge 86) |
| vercel:baseten | $0.0518 | $0.0271 | $0.1400 | $0.2034 | $0.1151 (judge 84) |
| direct:nebius | $0.0860 | $0.0969 | — | — | $0.1224 (judge 86) |
| direct:fireworks | $0.0713 | $0.0551 | $0.2280 | — | $0.1664 (judge 79) |
| openrouter:nebius/fp4 | $0.0604 | $0.0404 | $0.1712 | — | $0.1973 (judge 80) |
| openrouter:moonshotai/mxfp4 | $0.0962 | $0.0482 | $0.2201 | $0.3761 | $0.2104 (judge 88) |
| openrouter:fireworks/fast | $0.0863 | $0.1039 | $0.2552 | — | $0.2428 (judge 84) |
One collapsed card per session. Header carries a pass/fail glyph, the lane, task, attempt number (out of repeats), cost, and mini judge score. Expand to see the full judge verdict + reasoning, the four judge sub-scores (Action / Grounding / Verification / Efficiency, each 0-25), the verification exit code, and a tail of the child process's stdout.
tool_call. Deterministic.unique tool signatures / total tool calls. Penalizes duplicate calls.Sessions are sorted with passing high-judge-score first; failures and aborts sink to the bottom.
9cf8c41f-2795-47fe-b8ff-06551fee006805b8226fc-8f5a-4c5c-9e4b-25ce11a50d110ccacec46-4d87-43f9-acb9-63b7acfdf31c0b50cc8aa-ea44-4f47-a2c2-6adf632e89a10b5d562f9-368d-4550-9e39-0314ddd5f2f80bbefc9bd-5866-4ca9-8404-ee2ccae056ae0a1dd04d3-6184-4b8c-820f-e6559edf7148015819535-b93e-4691-8f93-148cf024c5130c8c5e140-399a-49dd-8df2-9d4b5719b7d409e05e236-c5be-43b5-a9c5-d83860bee68801fbbf727-a38c-4c97-9736-7cb912b4b1a70f38517b3-2aec-4cb8-b04b-da4eab70da8b042bd03ea-570e-467c-a3e0-fb081ddbdf030f0fd00ff-4282-46f0-90bb-37481030e1c50e1ab06f4-44a5-4538-bc19-eca9ab6a6360063e1a150-42b7-44fc-96f8-b63e8587c012051e01eb5-beef-4791-92f3-63f39eb4f5cc0a2f4303b-7a4a-457c-8130-37c5b3b385eb036bf20b9-cdae-436b-b2fb-a065448903b40bea2e855-1f57-42d1-bfa3-db1cc497085b04dad49e6-d3e7-4b64-8134-5ea0fcf8fcdd0392e75cf-489b-43b3-aa35-a24ccd669c0e0e41c7477-9c85-431a-8b57-f03642f7fef10ace2daad-8f9b-4079-880c-a8d54aa77b6a0e48718ac-948c-48b1-9501-d25d4f072f5f029ed1313-923a-4e84-952e-36e1d2702faf09e61bb6e-0c2a-4282-a147-83e25007ab4c0ff2cf1ea-b510-4d22-849f-4e5a7a3149ce00acb4722-8760-420a-b483-62eb16881e400dccbd7a2-94dd-47f2-9348-1f8d7f38bd790e3c065a9-1315-4918-81aa-bf59d2f6f86c05092a516-4370-4dd8-9450-9b9dddba58430b624460d-2541-4cef-8afe-f42847f98ae1063cbb486-2280-4fe8-9f4d-832cfcb7bb0f0349800a5-2768-446f-8f7f-f68019032aa1048b2164f-eb2d-4765-9b42-aef24edb485e0bcfc62a1-1be1-4412-a392-c784bfbe9f01098dc18d9-fbf6-4e40-a325-0868878d0496033fc2fef-8682-47af-b717-755184b9182c0178eb56d-c72a-4e78-9777-0dbb828c2c1808e18f46e-599a-428a-b096-911a751eae4f028443802-ce0a-426c-aca0-fc665a18564c0a2cd6456-84ee-4e6c-9714-3e003922a5ca0fe241c9c-723e-420d-87d7-743cc71835da08643bd7c-aa27-4bee-83d5-525db33390b50de743f72-5f70-4f7b-8609-ee23963744ee0c8aa45d7-fca7-4701-af4e-ce8610cc32450db8c1b24-fe07-4bca-be45-9c9644e73acd065a789c5-00b5-4308-94e7-9762c0bfa6a80b9911169-676d-439d-b53c-8ff85c02fa5e05aa48f01-60bf-4a4a-9a3c-a9d61c0f322d0573ad694-dfc5-4786-8b66-5c675ac1629a0a50cf5cd-5f22-4e13-bad3-9dd91db61aba0dadfdf41-bbb3-4911-a905-035f8058fdd90c6deb46d-741e-4769-9054-d56cfb0f63c70aae76f5c-8b1a-43c3-9bf1-b3d9208c35910f8c9acf3-0c0a-4b4c-ac07-0e51b16adf080ee8a3587-54cd-4cdd-a026-cfb3c1100db50019dfbd9-0adc-4f99-a760-414b5118bbbc078926d2e-c761-4326-a499-b8a18d84182c0570370e8-c58e-4532-b95f-bcbef209da210559d3f1f-c5e6-469f-ae48-63386b77fea90c710990b-29f2-4b88-a46d-2e08004fa5f70d6c4a1ef-5e91-4b12-b5ad-8bb2209dd146037e84eff-0f9a-46ce-ac12-2a79bfd5ee800aad4889d-f7f0-46ed-88fb-4eef6166e54d05f268b65-6326-4c02-ada0-6b3755ab6db10eba8256f-8f80-47fb-be14-66378e674cb3073f1197c-3cd3-4c92-80c7-da0a8440b8030123e3f21-23c2-45cd-8eca-b11a81c56f14060999f4b-e3e5-42f4-83ce-5c87a0ef2830060a3bf9e-7862-47b6-9053-dbfaa49c8db200e08a2f5-6177-4ac5-b693-8d57fd3ad7f908397ded7-88b1-473d-b142-d0a85fa8930d01fa740ca-01f7-4602-9b9e-2bb971574f4f0c8b29c92-b939-4a71-8df8-d2a312fb58a104514590f-0c45-4d53-a5e3-cc8384354540069b5302d-400a-4029-9289-ff4a3a0d4c0903a5d0acf-acdd-4fc2-915d-b099f6c867f901c26453b-ba09-4e39-b397-bc581f4fd48701ab94569-39b1-478b-a3bc-614ca4275b310f0051fd9-cd4c-4b83-bb6b-b2d77ba510180d84e54f0-d02c-4d39-af4e-3577dbd392b70371b5aec-3f4e-4ff1-ac00-9a2787bf3dc30947a52b6-79af-4077-9506-bea31fe8274e04c1954c4-7143-47b4-890a-42625bd72ba80c8196a5f-8d25-4ef5-bb6e-d63749a78adbnull (timeout)5069d63f-72bb-4be1-a8e7-4388bb1d211fnull (timeout)5302e087-552a-437f-b955-06c959f62f841ff22c858-77ce-4638-a2dc-67f0a825ba771a06d80ad-de41-4dcf-914b-f66a72715aa817d553a3b-4c2b-440c-ad34-d21696a64d601f9e2c483-2bbd-4e02-9b7e-1a5cf8cf390f1ee2cb854-0868-41a1-ac7e-cc66d4ca3c0a17c30bb2a-1c13-4f82-9747-2a2eb25779d510f12ddf6-6b4c-49ba-aaf5-200b8355af2e1433b7ff5-63ed-4392-89d0-151739f0516d17d66055a-5541-4ec2-a3dc-54d569048721null (timeout)7b14bf8e-535f-4470-8dac-9e9dee18bf211775f4272-ac8c-4235-a0ad-2faba440bf66null (timeout)130aa01d-b18d-42fc-9860-f23998608cc3null (timeout)ca4b0e16-07cb-44cb-a9f0-483125b46308nullc7916178-5a10-4e90-a056-cf4a36b0c880null81595e55-0433-475a-8dfc-e32f71eddc4bnullfd6737e5-234e-40b1-820e-71a2a487286anull9f1e324b-bacd-463d-813d-ff661ded6621nullf2195f10-5c05-42d4-bf6f-91129d268383null0de5baaf-afd6-4948-a0b6-a146cb7fbe96nulla34bbad3-b99a-4bc5-83d6-60605c03b20enull2640ee5e-8c52-41c8-ba1d-ea451dd31cb9null4d72a041-617f-43c6-a244-d41cff400911null7f54f4a7-b5e2-412b-b59f-a3eaacf95128nullc9e13a47-3e17-4dbf-96b1-44809a218c8fnull66e5939f-9288-4882-af0b-10e2e826eacdnull64e9d35e-3a00-4619-b6d3-4056efa47828null62980829-c4b6-4b3c-a72b-1e85d7a6c971nulle8efa75b-0592-444c-90c9-a2f20cfe3387nullbdf56186-38f8-4274-81c9-2784b54bad49null8961d1dd-3227-4208-943f-17f69808bad2nullafe5dc3b-a5d2-4e4b-b745-458e348615f5nulldirect:moonshot-intl) beats best OR route (openrouter:nebius/fp4) by 60.0% weighted.
Every field the runner recorded for this run, so we can decide which to promote into real report sections. Look at Layer 2 first — dead fields (single-value columns) are dimmed and marked; signal fields (spread > 1 unique) are candidates for the leaderboard, lane table, and drilldown. Layer 3 opens up the JSON bodies where the interesting content actually lives.
Rip this section out once phase 1 lands.
| Column | Distribution | Sample |
|---|---|---|
| id | 1 unique (1 rows) | |
| started_at | 20:05:06 → 20:05:06 (span 0s) | |
| finished_at | 21:12:17 → 21:12:17 (span 0s) | |
| status | complete 1 | |
| target_dir | 1 rows · avg length 50 chars | /Users/rupulsafaya/Documents/GitHub/canaryone-demo |
| meta_json | 1 rows · avg length 1317 chars | {"runId":"e860167a-2cfb-4415-a19c-86b476b2a1d6","startedAt":"2026-07-29T20:05:06.013Z","targetDir":"… |
| Column | Distribution | Sample |
|---|---|---|
| id | 120 unique (120 rows) | |
| run_id | 1 unique (120 rows) | |
| task_id | 4 unique (120 rows) | |
| task_file | 120 rows · avg length 23 chars | tests/t3-difficult.test.js |
| model_slug | moonshotai/kimi-k3 84accounts/fireworks/routers/kimi-k3-fast 12kimi-k3 12moonshotai/Kimi-K3 12 | |
| destination_slug | direct:fireworks 12direct:moonshot-intl 12direct:nebius 12vercel:baseten 12vercel:fireworks 12 +5 more | |
| router | vercel 48direct 36openrouter 36 | |
| repeat_ix | min 0 · max 2 · avg 1 |
|
| status | complete 86queued 19failed 15 | |
| started_at | 20:05:06 → 21:08:28 (span 63m22s) | |
| finished_at | 20:05:54 → 21:10:12 (span 64m17s) | |
| cost_usd | min 0 · max 0.4195 · avg 0.0877 |
|
| verify_exit_code | min 0 · max 1 · avg 0.1042 (24 null) |
|
| verify_stdout_tail | 101 rows · avg length 559 chars (19 null) | (node:93894) ExperimentalWarning: SQLite is an experimental feature and might change at any time (Us… |
| verify_stderr_tail | 101 rows · avg length 0 chars (19 null) | |
| failure_class | timeout 5 | |
| worktree_path | 101 rows · avg length 138 chars (19 null) | /Users/rupulsafaya/Documents/GitHub/canaryone-demo/.c1/worktrees/e860167a-2cfb-4415-a19c-86b476b2a1d… |
| proxy_port | min 60767 · max 61998 · avg 61246.9901 (19 null) |
| Column | Distribution | Sample |
|---|---|---|
| id | 705 unique (705 rows) | |
| session_id | 101 unique (705 rows) | |
| step_ix | min 0 · max 14 · avg 3.4113 |
|
| started_at | 20:05:06 → 21:10:02 (span 64m55s) | |
| finished_at | 20:05:10 → 21:10:12 (span 65m01s) | |
| http_status | min 200 · max 412 · avg 200.6110 (11 null) |
|
| inbound_shape | openai 705 | |
| path | /v1/chat/completions 705 | |
| input_tokens | min 0 · max 14700 · avg 3307.4723 |
|
| output_tokens | min 0 · max 1500 · avg 294.6241 |
|
| cost_usd | min 0 · max 0.0714 · avg 0.0154 |
|
| latency_ms | min 44 · max 303932 · avg 16353.0156 |
|
| translation_notes | — (all null, 705 rows) | |
| traffic_log_offset | min 174 · max 12902010 · avg 5399357.4851 |
|
| traffic_log_length | min 3240 · max 872382 · avg 53157.5319 |
|
| failure_class | forward_failed 11http_412 2 |
| Column | Distribution | Sample |
|---|---|---|
| id | 808 unique (808 rows) | |
| session_id | 101 unique (808 rows) | |
| dimension | action_score 101efficiency_score 101grounding_score 101judge_reasoning 101outcome 101 +3 more | |
| value | 25 176success 8820 8821 3618 25 +234 more | |
| confidence | min 0.5000 · max 1 · avg 0.9292 |
|
| generated_at | 21:10:16 → 21:12:17 (span 2m01s) | |
| model | anthropic/claude-haiku-4.5 808 | |
| classifier_id | canaryone_judge_v1_local 808 | |
| classifier_version | 2026-07-29-haiku-r5-local 808 |
| Column | Distribution | Sample |
|---|---|---|
| task_id | 5 unique (5 rows) | |
| file | tests/t1-easy.test.js 1tests/t2-medium.test.js 1tests/t3-difficult.test.js 1tests/t4-super.test.js 1tests/q5-bramblegate.test.js 1 | |
| summary | 5 rows · avg length 93 chars | Agent autonomously diagnoses and fixes a buggy SQL query in churn.js via read/run/write/verify loop. |
| uses_llm | min 1 · max 1 · avg 1 |
. object
.max_tokens number
min 1500, max 1500, avg 1500
.messages array
length 2..39 (avg 11.7)
.messages[] object
.messages[].annotations null
present 267/701 (38%)
.messages[].audio null
present 267/701 (38%)
.messages[].content string|null
len 0..11859 (avg 732)
.messages[].function_call null
present 267/701 (38%)
.messages[].provider_metadata object
.messages[].provider_metadata.baseten object
present 349/701 (50%)
.messages[].provider_metadata.baseten.acceptedPredictionTokens number
present 349/701 (50%) · min 21, max 254, avg 66.4556
.messages[].provider_metadata.baseten.rejectedPredictionTokens number
present 349/701 (50%) · min 0, max 0, avg 0
.messages[].provider_metadata.fireworks object
present 196/701 (28%)
.messages[].provider_metadata.gateway object
.messages[].provider_metadata.gateway.cost string
len 6..9 (avg 8)
.messages[].provider_metadata.gateway.gatewayCost string
len 6..9 (avg 8)
.messages[].provider_metadata.gateway.generationId string
len 30..30 (avg 30)
.messages[].provider_metadata.gateway.inferenceCost string
len 6..9 (avg 8)
.messages[].provider_metadata.gateway.inputInferenceCost string
len 6..9 (avg 8)
.messages[].provider_metadata.gateway.marketCost string
len 6..9 (avg 8)
.messages[].provider_metadata.gateway.outputInferenceCost string
len 5..8 (avg 8)
.messages[].provider_metadata.gateway.routing object
.messages[].provider_metadata.gateway.routing.canonicalSlug string
len 18..18 (avg 18) · moonshotai/kimi-k3 813
.messages[].provider_metadata.gateway.routing.fallbacksAvailable array
length 0..0 (avg 0.0)
.messages[].provider_metadata.gateway.routing.finalProvider string
len 6..10 (avg 8) · baseten 349fireworks 196moonshotai 201nebius 67
.messages[].provider_metadata.gateway.routing.modelAttemptCount number
min 1, max 1, avg 1
.messages[].provider_metadata.gateway.routing.modelAttempts array
length 1..1 (avg 1.0)
.messages[].provider_metadata.gateway.routing.modelAttempts[] object
.messages[].provider_metadata.gateway.routing.modelAttempts[].canonicalSlug string
len 18..18 (avg 18) · moonshotai/kimi-k3 813
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttemptCount number
min 1, max 1, avg 1
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts array
length 1..1 (avg 1.0)
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts[] object
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].credentialType string
len 6..6 (avg 6) · system 813
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].endTime number
min 1785355777277, max 1785359402023, avg 1785356857791.6714
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].provider string
len 6..10 (avg 8) · baseten 349fireworks 196moonshotai 201nebius 67
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].providerRequestId string
present 587/701 (84%) · len 32..41 (avg 37)
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].providerResponseId string
len 32..41 (avg 38)
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].startTime number
min 1785355764070, max 1785359399406, avg 1785356847410.5623
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].statusCode number
min 200, max 200, avg 200
.messages[].provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].success boolean
.messages[].provider_metadata.gateway.routing.modelAttempts[].success boolean
.messages[].provider_metadata.gateway.routing.originalModelId string
len 18..18 (avg 18) · moonshotai/kimi-k3 813
.messages[].provider_metadata.gateway.routing.planningReasoning string
len 113..125 (avg 119)
.messages[].provider_metadata.gateway.routing.resolvedProvider string
len 6..10 (avg 8) · baseten 349fireworks 196moonshotai 201nebius 67
.messages[].provider_metadata.gateway.routing.totalProviderAttemptCount number
min 1, max 1, avg 1
.messages[].provider_metadata.gateway.surchargeCost string
len 1..1 (avg 1) · 0 813
.messages[].provider_metadata.moonshotai object
present 201/701 (29%)
.messages[].provider_metadata.nebius object
present 67/701 (10%)
.messages[].reasoning string|null
len 31..4644 (avg 511) · Let me explore the directories. 8Let me look at the directory structure. 4Let me read the docs and schema first. 5Let me read the docs and README first. 5Let me use ls recursively instead. 5
.messages[].reasoning_content string|null
present 559/701 (80%) · len 0..4211 (avg 377) · Let me explore the directories. 2Let me look at the directory structure. 4Let me list directories. 4 11
.messages[].reasoning_details array
length 1..1 (avg 1.0)
.messages[].reasoning_details[] object
.messages[].reasoning_details[].format string
len 7..7 (avg 7) · unknown 1328
.messages[].reasoning_details[].index number
min 0, max 0, avg 0
.messages[].reasoning_details[].text string
len 31..4644 (avg 523) · Let me explore the directories. 8Let me look at the directory structure. 4Let me read the docs and schema first. 5Let me read the docs and README first. 5Let me use ls recursively instead. 5
.messages[].reasoning_details[].type string
len 14..14 (avg 14) · reasoning.text 1328
.messages[].refusal null
.messages[].role string
len 4..9 (avg 6) · system 701user 701assistant 2374tool 4446
.messages[].tool_call_id string
len 11..29 (avg 12)
.messages[].tool_calls array
length 1..6 (avg 1.9)
.messages[].tool_calls[] object
.messages[].tool_calls[].function object
.messages[].tool_calls[].function.arguments string
len 12..2192 (avg 75)
.messages[].tool_calls[].function.name string
len 9..12 (avg 10) · run_shell 1944read_file 1483sqlite_query 999write_file 20
.messages[].tool_calls[].id string
len 11..29 (avg 12)
.messages[].tool_calls[].index number
min 0, max 5, avg 0.6568
.messages[].tool_calls[].name null
present 279/701 (40%)
.messages[].tool_calls[].type string
len 8..8 (avg 8) · function 4446
.messages[].tools null
present 140/701 (20%)
.model string
len 11..11 (avg 11) · gpt-4o-mini 701
.tool_choice string
len 4..4 (avg 4) · auto 701
.tools array
length 4..4 (avg 4.0)
.tools[] object
.tools[].function object
.tools[].function.description string
len 108..169 (avg 140)
.tools[].function.name string
len 9..12 (avg 10) · read_file 701write_file 701sqlite_query 701run_shell 701
.tools[].function.parameters object
.tools[].function.parameters.properties object
.tools[].function.parameters.properties.cmd object
.tools[].function.parameters.properties.cmd.description string
len 70..70 (avg 70)
.tools[].function.parameters.properties.cmd.type string
len 6..6 (avg 6) · string 701
.tools[].function.parameters.properties.content object
.tools[].function.parameters.properties.content.description string
len 71..71 (avg 71)
.tools[].function.parameters.properties.content.type string
len 6..6 (avg 6) · string 701
.tools[].function.parameters.properties.db object
.tools[].function.parameters.properties.db.description string
len 60..60 (avg 60)
.tools[].function.parameters.properties.db.type string
len 6..6 (avg 6) · string 701
.tools[].function.parameters.properties.path object
.tools[].function.parameters.properties.path.description string
len 59..61 (avg 60)
.tools[].function.parameters.properties.path.type string
len 6..6 (avg 6) · string 1402
.tools[].function.parameters.properties.sql object
.tools[].function.parameters.properties.sql.description string
len 53..53 (avg 53)
.tools[].function.parameters.properties.sql.type string
len 6..6 (avg 6) · string 701
.tools[].function.parameters.required array
length 1..2 (avg 1.5)
.tools[].function.parameters.required[] string
len 2..7 (avg 4) · path 1402content 701db 701sql 701cmd 701
.tools[].function.parameters.type string
len 6..6 (avg 6) · object 2804
.tools[].type string
len 8..8 (avg 8) · function 2804
. object
.choices array
present 689/691 (100%) · length 1..1 (avg 1.0)
.choices[] object
present 689/691 (100%)
.choices[].finish_reason string
present 689/691 (100%) · len 4..10 (avg 9) · tool_calls 591stop 88length 10
.choices[].index number
present 689/691 (100%) · min 0, max 0, avg 0
.choices[].logprobs null
present 587/691 (85%)
.choices[].matched_stop null|number
present 69/691 (10%) · min 163586, max 163586, avg 35205.1683
.choices[].message object
present 689/691 (100%)
.choices[].message.annotations null
present 69/691 (10%)
.choices[].message.audio null
present 69/691 (10%)
.choices[].message.content null|string
present 689/691 (100%) · len 0..2232 (avg 329)
.choices[].message.function_call null
present 69/691 (10%)
.choices[].message.provider_metadata object
present 246/691 (36%)
.choices[].message.provider_metadata.baseten object
present 88/691 (13%)
.choices[].message.provider_metadata.baseten.acceptedPredictionTokens number
present 88/691 (13%) · min 20, max 258, avg 76.4205
.choices[].message.provider_metadata.baseten.rejectedPredictionTokens number
present 88/691 (13%) · min 0, max 0, avg 0
.choices[].message.provider_metadata.fireworks object
present 62/691 (9%)
.choices[].message.provider_metadata.gateway object
present 246/691 (36%)
.choices[].message.provider_metadata.gateway.cost string
present 246/691 (36%) · len 6..9 (avg 8)
.choices[].message.provider_metadata.gateway.gatewayCost string
present 246/691 (36%) · len 6..9 (avg 8)
.choices[].message.provider_metadata.gateway.generationId string
present 246/691 (36%) · len 30..30 (avg 30)
.choices[].message.provider_metadata.gateway.inferenceCost string
present 246/691 (36%) · len 6..9 (avg 8)
.choices[].message.provider_metadata.gateway.inputInferenceCost string
present 246/691 (36%) · len 6..9 (avg 8)
.choices[].message.provider_metadata.gateway.marketCost string
present 246/691 (36%) · len 6..9 (avg 8)
.choices[].message.provider_metadata.gateway.outputInferenceCost string
present 246/691 (36%) · len 5..8 (avg 7)
.choices[].message.provider_metadata.gateway.routing object
present 246/691 (36%)
.choices[].message.provider_metadata.gateway.routing.canonicalSlug string
present 246/691 (36%) · len 18..18 (avg 18) · moonshotai/kimi-k3 246
.choices[].message.provider_metadata.gateway.routing.fallbacksAvailable array
present 246/691 (36%) · length 0..0 (avg 0.0)
.choices[].message.provider_metadata.gateway.routing.finalProvider string
present 246/691 (36%) · len 6..10 (avg 8) · baseten 88fireworks 62moonshotai 63nebius 33
.choices[].message.provider_metadata.gateway.routing.modelAttemptCount number
present 246/691 (36%) · min 1, max 1, avg 1
.choices[].message.provider_metadata.gateway.routing.modelAttempts array
present 246/691 (36%) · length 1..1 (avg 1.0)
.choices[].message.provider_metadata.gateway.routing.modelAttempts[] object
present 246/691 (36%)
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].canonicalSlug string
present 246/691 (36%) · len 18..18 (avg 18) · moonshotai/kimi-k3 246
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttemptCount number
present 246/691 (36%) · min 1, max 1, avg 1
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts array
present 246/691 (36%) · length 1..1 (avg 1.0)
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts[] object
present 246/691 (36%)
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].credentialType string
present 246/691 (36%) · len 6..6 (avg 6) · system 246
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].endTime number
present 246/691 (36%) · min 1785355777277, max 1785359412312, avg 1785356675326.2234
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].provider string
present 246/691 (36%) · len 6..10 (avg 8) · baseten 88fireworks 62moonshotai 63nebius 33
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].providerRequestId string
present 181/691 (26%) · len 32..41 (avg 37)
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].providerResponseId string
present 246/691 (36%) · len 32..41 (avg 38)
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].startTime number
present 246/691 (36%) · min 1785355764070, max 1785359402226, avg 1785356663086.0125
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].statusCode number
present 246/691 (36%) · min 200, max 200, avg 200
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].providerAttempts[].success boolean
present 246/691 (36%)
.choices[].message.provider_metadata.gateway.routing.modelAttempts[].success boolean
present 246/691 (36%)
.choices[].message.provider_metadata.gateway.routing.originalModelId string
present 246/691 (36%) · len 18..18 (avg 18) · moonshotai/kimi-k3 246
.choices[].message.provider_metadata.gateway.routing.planningReasoning string
present 246/691 (36%) · len 113..125 (avg 119)
.choices[].message.provider_metadata.gateway.routing.resolvedProvider string
present 246/691 (36%) · len 6..10 (avg 8) · baseten 88fireworks 62moonshotai 63nebius 33
.choices[].message.provider_metadata.gateway.routing.totalProviderAttemptCount number
present 246/691 (36%) · min 1, max 1, avg 1
.choices[].message.provider_metadata.gateway.surchargeCost string
present 246/691 (36%) · len 1..1 (avg 1) · 0 246
.choices[].message.provider_metadata.moonshotai object
present 63/691 (9%)
.choices[].message.provider_metadata.nebius object
present 33/691 (5%)
.choices[].message.reasoning string|null
present 419/691 (61%) · len 31..5334 (avg 698) · Let me explore the directories. 2Let me look at the directory structure. 1Let me read the docs and schema first. 1Let me read the docs and README first. 1Let me use ls recursively instead. 1
.choices[].message.reasoning_content string|null
present 170/691 (25%) · len 0..4597 (avg 556) · Let me explore the directories. 1Let me look at the directory structure. 1Let me list directories. 1 2
.choices[].message.reasoning_details array
present 400/691 (58%) · length 1..1 (avg 1.0)
.choices[].message.reasoning_details[] object
present 400/691 (58%)
.choices[].message.reasoning_details[].format string
present 400/691 (58%) · len 7..7 (avg 7) · unknown 400
.choices[].message.reasoning_details[].index number
present 400/691 (58%) · min 0, max 0, avg 0
.choices[].message.reasoning_details[].text string
present 400/691 (58%) · len 31..5334 (avg 709) · Let me explore the directories. 2Let me look at the directory structure. 1Let me read the docs and schema first. 1Let me read the docs and README first. 1Let me use ls recursively instead. 1
.choices[].message.reasoning_details[].type string
present 400/691 (58%) · len 14..14 (avg 14) · reasoning.text 400
.choices[].message.refusal null
present 341/691 (49%)
.choices[].message.role string
present 689/691 (100%) · len 9..9 (avg 9) · assistant 689
.choices[].message.tool_calls array|null
present 601/691 (87%) · length 0..6 (avg 1.9)
.choices[].message.tool_calls[] object
.choices[].message.tool_calls[].function object
.choices[].message.tool_calls[].function.arguments string
len 12..2192 (avg 95)
.choices[].message.tool_calls[].function.name string
len 9..12 (avg 10) · run_shell 432read_file 384sqlite_query 300write_file 6
.choices[].message.tool_calls[].id string
len 11..29 (avg 12)
.choices[].message.tool_calls[].index number
min 0, max 5, avg 0.6712
.choices[].message.tool_calls[].name null
present 75/691 (11%)
.choices[].message.tool_calls[].type string
len 8..8 (avg 8) · function 1122
.choices[].message.tools null
present 46/691 (7%)
.choices[].native_finish_reason string
present 272/691 (39%) · len 4..10 (avg 9) · tool_calls 233stop 29length 10
.created number
present 689/691 (100%) · min 1785355506, max 1785359412, avg 1785356946.3440
.error object
present 2/691 (0%)
.error.code string
present 2/691 (0%) · len 19..19 (avg 19) · PRECONDITION_FAILED 2
.error.message string
present 2/691 (0%) · len 193..193 (avg 193)
.error.param null
present 2/691 (0%)
.error.type string
present 2/691 (0%) · len 5..5 (avg 5) · error 2
.generationId string
present 246/691 (36%) · len 30..30 (avg 30)
.id string
present 689/691 (100%) · len 30..41 (avg 33)
.metadata object
present 69/691 (10%)
.metadata.weight_version string
present 69/691 (10%) · len 7..7 (avg 7) · default 69
.model string
present 689/691 (100%) · len 7..33 (avg 18) · moonshotai/kimi-k3 518accounts/fireworks/models/kimi-k3 46kimi-k3 56moonshotai/Kimi-K3 69
.moderation null
present 69/691 (10%)
.object string
present 689/691 (100%) · len 15..15 (avg 15) · chat.completion 689
.provider string
present 272/691 (39%) · len 5..12 (avg 8) · Moonshot AI 107BaseTen 3DigitalOcean 14Fireworks 7Nebius 82Together 24Modal 35
.request_id string
present 2/691 (0%) · len 41..41 (avg 41)
.service_tier null
present 341/691 (49%)
.system_fingerprint string|null
present 587/691 (85%) · len 12..13 (avg 13)
.usage object
present 689/691 (100%)
.usage.cache_creation_input_tokens number
present 246/691 (36%) · min 0, max 0, avg 0
.usage.cached_tokens number
present 56/691 (8%) · min 512, max 8192, avg 2016
.usage.completion_tokens number
present 689/691 (100%) · min 42, max 1500, avg 300.9042
.usage.completion_tokens_details object|null
present 643/691 (93%)
.usage.completion_tokens_details.audio_tokens number
present 272/691 (39%) · min 0, max 0, avg 0
.usage.completion_tokens_details.image_tokens number
present 518/691 (75%) · min 0, max 0, avg 0
.usage.completion_tokens_details.reasoning_tokens number
present 574/691 (83%) · min 0, max 1356, avg 140.8118
.usage.completion_tokens_details.reasoning_tokens_estimated boolean
present 87/691 (13%)
.usage.cost number
present 518/691 (75%) · min 0.0016, max 0.0627, avg 0.0112
.usage.cost_details object
present 518/691 (75%)
.usage.cost_details.upstream_inference_completions_cost number
present 518/691 (75%) · min 0, max 0.0225, avg 0.0029
.usage.cost_details.upstream_inference_cost number|null
present 518/691 (75%) · min 0.0016, max 0.0627, avg 0.0115
.usage.cost_details.upstream_inference_prompt_cost number
present 518/691 (75%) · min 0, max 0.0402, avg 0.0041
.usage.gateway_cost number
present 246/691 (36%) · min 0.0016, max 0.0352, avg 0.0090
.usage.is_byok boolean
present 518/691 (75%)
.usage.market_cost number
present 246/691 (36%) · min 0.0016, max 0.0352, avg 0.0090
.usage.prompt_tokens number
present 689/691 (100%) · min 677, max 14700, avg 3380.6705
.usage.prompt_tokens_details object|null
present 689/691 (100%)
.usage.prompt_tokens_details.audio_tokens number
present 518/691 (75%) · min 0, max 0, avg 0
.usage.prompt_tokens_details.cache_write_tokens number
present 272/691 (39%) · min 0, max 0, avg 0
.usage.prompt_tokens_details.cached_tokens number
present 620/691 (90%) · min 0, max 10752, avg 1502.5032
.usage.prompt_tokens_details.video_tokens number
present 518/691 (75%) · min 0, max 0, avg 0
.usage.reasoning_tokens number
present 69/691 (10%) · min 0, max 0, avg 0
.usage.total_tokens number
present 689/691 (100%) · min 736, max 15440, avg 3681.5747
. object
.configDir string
len 54..54 (avg 54)
.lanes array
length 10..10 (avg 10.0)
.lanes[] object
.lanes[].destination string
len 13..13 (avg 13)
.lanes[].model string
len 18..18 (avg 18)
.lanes[].router string
len 6..6 (avg 6)
.parallelism number
min 3, max 3, avg 3
.repeats number
min 3, max 3, avg 3
.runId string
len 36..36 (avg 36)
.startedAt string
len 24..24 (avg 24)
.targetDir string
len 50..50 (avg 50)
.tasks array
length 4..4 (avg 4.0)
.tasks[] object
.tasks[].file string
len 22..22 (avg 22)
.tasks[].id string
len 3..3 (avg 3)
src/runner/print-summary.ts)| Total spend | $10.53 |
| Pass rate | 86/120 |
| Judge score range | 25–93 |
| Best value | vercel:nebius $0.0686 weighted/pass |
| Cheapest raw | vercel:nebius $0.0472/pass (traj 69) |