Default cost is real cost
Every tool schema, skill description and system instruction is sent on the very first turn. That baseline is paid on each request, before any useful work happens.
hibench measures the default context footprint of coding agents: system prompts, tools, skills, MCP servers, and sub-agents loaded before the first user request is answered.
Benchmark data last updated July 7, 2026 UTC
Every tool schema, skill description and system instruction is sent on the very first turn. That baseline is paid on each request, before any useful work happens.
We capture the first real outbound request and count it with a single fixed tokenizer so numbers stay comparable across agents, versions and models.
Footprints drift over releases. hibench keeps one canonical capture per version so you can watch context grow (or shrink) over time.
Isolate
Run each agent in Docker inside a fresh, empty Git repo.
Intercept
Point it at a local recorder with a dummy key — no upstream model call.
Capture
Send the prompt Hi and record the first outbound request body.
Count
Tokenize every field with o200k_base and break it down.
20,556
o200k_base total tokens · 25 tools · 116 versions
18,863
o200k_base total tokens · 33 tools · 67 versions
16,390
o200k_base total tokens · 15 tools · 9 versions
13,062
o200k_base total tokens · 15 tools · 61 versions
13,023
o200k_base total tokens · 10 tools · 115 versions

12,223
o200k_base total tokens · 27 tools · 8 versions
11,868
o200k_base total tokens · 7 tools · 9 versions

11,713
o200k_base total tokens · 23 tools · 3 versions

9,426
o200k_base total tokens · 13 tools · 111 versions
9,287
o200k_base total tokens · 24 tools · 106 versions
8,710
o200k_base total tokens · 11 tools · 116 versions

8,513
o200k_base total tokens · 8 tools · 101 versions

6,671
o200k_base total tokens · 9 tools · 109 versions
5,785
o200k_base total tokens · 9 tools · 56 versions

4,530
o200k_base total tokens · 25 tools · 36 versions
1,264
o200k_base total tokens · 4 tools · 109 versions