|
| 1 | +# Benchmark Report |
| 2 | + |
| 3 | +- Suite: `release` |
| 4 | +- Started: `2026-04-11 20:43:07 MSK` |
| 5 | +- Finished: `2026-04-11 20:48:13 MSK` |
| 6 | +- Environment: `windows/amd64` |
| 7 | +- Model evaluation: enabled (`z-ai/glm-5.1` via `https://openrouter.ai/api/v1`) |
| 8 | + |
| 9 | +Task success rate in this report means: correct tool selection + schema-valid arguments + local task oracle assertions. It is not a claim of end-to-end external API execution success. |
| 10 | + |
| 11 | +## Scenario Summary |
| 12 | + |
| 13 | +| Scenario | Layer | Tools | Namespaces | Compile | Index | Namespace | Schema | Invoke | |
| 14 | +| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| 15 | +| synthetic-10 | synthetic | 10 | 10 | 517.000 us | 5.144 us | 5.174 us | 0 | 0 | |
| 16 | +| synthetic-50 | synthetic | 50 | 10 | 2.250 ms | 29.942 us | 0 | 0 | 0 | |
| 17 | +| synthetic-200 | synthetic | 200 | 10 | 6.433 ms | 89.988 us | 5.766 us | 0 | 0 | |
| 18 | +| synthetic-500 | synthetic | 500 | 10 | 15.539 ms | 145.770 us | 0 | 8.842 us | 7.331 us | |
| 19 | +| bfcl-live-simple-subset | real-world | 4 | 1 | 0 | 4.630 us | 0 | 0 | 0 | |
| 20 | +| github-collaboration-subset | real-world | 14 | 2 | 4.570 ms | 9.985 us | 0 | 65.863 us | 0 | |
| 21 | +| github-rest-full-metrics | real-world | 1112 | 44 | 118.652 ms | 218.875 us | 9.994 us | 0 | 0 | |
| 22 | + |
| 23 | +## synthetic-10 |
| 24 | + |
| 25 | +Small synthetic catalog with model evaluation enabled. |
| 26 | + |
| 27 | +| Variant | Avg visible tools | Avg token proxy | Wrong-tool rate | Invalid-args rate | Task success rate | Avg latency | Useful selection | |
| 28 | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| 29 | +| raw | 10.0 | 910.0 | 0.00 | 0.00 | 1.00 | 1.259 s | 1.259 s | |
| 30 | +| baseline | 2.3 | 246.3 | 0.00 | 0.00 | 1.00 | 750.658 ms | 750.658 ms | |
| 31 | +| compiled | 10.0 | 448.0 | 0.00 | 0.00 | 1.00 | 662.831 ms | 662.831 ms | |
| 32 | + |
| 33 | + |
| 34 | +## synthetic-50 |
| 35 | + |
| 36 | +Medium synthetic catalog for structural metrics only in the release run. |
| 37 | + |
| 38 | +| Variant | Avg visible tools | Avg token proxy | Wrong-tool rate | Invalid-args rate | Task success rate | Avg latency | Useful selection | |
| 39 | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| 40 | +| raw | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 41 | + |
| 42 | +Skipped reason for `raw`: model evaluation disabled for scenario |
| 43 | + |
| 44 | +| baseline | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 45 | + |
| 46 | +Skipped reason for `baseline`: model evaluation disabled for scenario |
| 47 | + |
| 48 | +| compiled | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 49 | + |
| 50 | +Skipped reason for `compiled`: model evaluation disabled for scenario |
| 51 | + |
| 52 | + |
| 53 | + |
| 54 | +## synthetic-200 |
| 55 | + |
| 56 | +Large synthetic catalog for structural metrics only in the release run. |
| 57 | + |
| 58 | +| Variant | Avg visible tools | Avg token proxy | Wrong-tool rate | Invalid-args rate | Task success rate | Avg latency | Useful selection | |
| 59 | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| 60 | +| raw | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 61 | + |
| 62 | +Skipped reason for `raw`: model evaluation disabled for scenario |
| 63 | + |
| 64 | +| baseline | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 65 | + |
| 66 | +Skipped reason for `baseline`: model evaluation disabled for scenario |
| 67 | + |
| 68 | +| compiled | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 69 | + |
| 70 | +Skipped reason for `compiled`: model evaluation disabled for scenario |
| 71 | + |
| 72 | + |
| 73 | + |
| 74 | +## synthetic-500 |
| 75 | + |
| 76 | +Very large synthetic catalog for structural metrics only in the release run. |
| 77 | + |
| 78 | +| Variant | Avg visible tools | Avg token proxy | Wrong-tool rate | Invalid-args rate | Task success rate | Avg latency | Useful selection | |
| 79 | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| 80 | +| raw | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 81 | + |
| 82 | +Skipped reason for `raw`: model evaluation disabled for scenario |
| 83 | + |
| 84 | +| baseline | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 85 | + |
| 86 | +Skipped reason for `baseline`: model evaluation disabled for scenario |
| 87 | + |
| 88 | +| compiled | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 89 | + |
| 90 | +Skipped reason for `compiled`: model evaluation disabled for scenario |
| 91 | + |
| 92 | + |
| 93 | + |
| 94 | +## bfcl-live-simple-subset |
| 95 | + |
| 96 | +BFCL subset with model evaluation enabled for a bounded set of tasks. |
| 97 | + |
| 98 | +| Variant | Avg visible tools | Avg token proxy | Wrong-tool rate | Invalid-args rate | Task success rate | Avg latency | Useful selection | |
| 99 | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| 100 | +| raw | 4.0 | 742.0 | 0.00 | 0.00 | 1.00 | 3.203 s | 3.203 s | |
| 101 | +| baseline | 1.8 | 362.8 | 0.00 | 0.00 | 1.00 | 921.339 ms | 921.339 ms | |
| 102 | +| compiled | 4.0 | 188.0 | 0.25 | 0.50 | 0.25 | 3.182 s | 10.067 s | |
| 103 | + |
| 104 | + |
| 105 | +## github-collaboration-subset |
| 106 | + |
| 107 | +Official GitHub REST API subset with model evaluation enabled for a bounded set of tasks. |
| 108 | + |
| 109 | +| Variant | Avg visible tools | Avg token proxy | Wrong-tool rate | Invalid-args rate | Task success rate | Avg latency | Useful selection | |
| 110 | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| 111 | +| raw | 14.0 | 6322.0 | 0.25 | 0.50 | 0.00 | 4.949 s | 0 | |
| 112 | +| baseline | 5.0 | 2576.8 | 0.25 | 0.25 | 0.25 | 3.062 s | 1.227 s | |
| 113 | +| compiled | 14.0 | 360.0 | 0.50 | 0.50 | 0.50 | 3.005 s | 4.482 s | |
| 114 | + |
| 115 | + |
| 116 | +## github-rest-full-metrics |
| 117 | + |
| 118 | +Full official GitHub REST API description for import/compile/gateway metrics only. |
| 119 | + |
| 120 | +| Variant | Avg visible tools | Avg token proxy | Wrong-tool rate | Invalid-args rate | Task success rate | Avg latency | Useful selection | |
| 121 | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| 122 | +| raw | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 123 | + |
| 124 | +Skipped reason for `raw`: scenario has no tasks |
| 125 | + |
| 126 | +| baseline | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 127 | + |
| 128 | +Skipped reason for `baseline`: scenario has no tasks |
| 129 | + |
| 130 | +| compiled | 0.0 | 0.0 | 0.00 | 0.00 | 0.00 | 0 | 0 | |
| 131 | + |
| 132 | +Skipped reason for `compiled`: scenario has no tasks |
| 133 | + |
| 134 | + |
| 135 | +Notes: |
| 136 | +- scenario contains no task oracle; only structural metrics, compile time, and gateway overhead were measured |
| 137 | + |
| 138 | + |
0 commit comments