claude-sonnet scored higher on quality, but gemini-pro bid cheaper and responds faster — enough to come out ahead with priorities set to quality 75, cost 15, speed 10.
utility = 0.75 × quality + 0.15 × cost + 0.1 × speed − uncertainty penalty
| Bid | Quality (× 0.75) | Cost (× 0.15) | Speed (× 0.1) | Penalty | Utility |
|---|---|---|---|---|---|
| gemini-pro | 0.78 → 0.585 | 0.80 → 0.120 | 0.70 → 0.070 | 0 | 0.775 |
| claude-sonnet | 0.82 → 0.615 | 0.55 → 0.083 | 0.60 → 0.060 | 0 | 0.758 |
| gpt-mini | 0.74 → 0.555 | 0.90 → 0.135 | 0.40 → 0.040 | 0 | 0.730 |
Quality is a normalized composite of public benchmark results — a routing proxy, not a success probability. Cost and speed are scored against fixed bounds saved with this run, so one bid never shifts another’s score.
| Model · deployment | Reason | What would change it |
|---|---|---|
| qwen-coder-q4 · operator vLLM | Quantized build; its benchmark scores are inherited from full-precision weights, so they count only as a proxy | Turn on “allow missing coverage” — joins with an uncertainty penalty |
| small-flash · Google | Coding quality 0.51, below the policy minimum of 0.60 | Lower the minimum quality in the policy |
| legacy-chat · OpenAI-compatible | Endpoint probe found no tool-call support; this request declares a tool | Hard requirement — can’t be relaxed |
| Deployment | Arrived | In / out per Mtok | Strategy | Queue delay | Status |
|---|---|---|---|---|---|
| gpt-mini | 84 ms | — / — | fixed rate | — | valid |
| claude-sonnet | 142 ms | — / — | capacity-adjusted | — | valid |
| gemini-pro | 205 ms | — / — | bounded discount | — | valid · awarded |
| open-coder | 341 ms | — / — | fixed rate | — | rejected · arrived after deadline |
Bids were sealed. Invitations carried only an auction ID, each bidder’s eligible offering IDs, the deadline and protocol version — no prompt, classification or app weights. Rates are simulated demo offers, not vendor discounts.
Estimates can differ from the provider’s bill. Figures are kept separate rather than blended.