We study how information representation affects reasoning quality, context requirements, cost and reliability.
How much usable meaning a representation carries per unit of size — and how that changes what a model can do with it.
How much correct reasoning a model performs per unit of context it is given.
The distance between human intent and the information a machine is given. The larger it is, the more reasoning is spent reconstructing what was already meant.
How much of the context supplied to a model actually contributes to the answer.
How the same model's accuracy, cost and latency move when only the intelligence it receives changes.
What one completed customer outcome costs across a production agent workload, rather than what one call costs.
The industry measures models. We measure what a model is given. These are the units we report in, defined tightly enough to be computed the same way twice.
Useful units of meaning per KB. How much a representation says in the space it takes.
Correct task outcomes per unit of context. What a token of context is actually worth.
Distance between human intent and machine-readable information. How far a system has to travel before it can start reasoning.
FROM BENCHMARK #001. EVERY FIGURE WE PUBLISH CARRIES ITS METHOD, SCOPE AND LIMITATIONS.
What it is for, what would falsify it, and where it currently stands.
A task-conditioned unit for the work that surrounds code generation — the load an agentic system actually carries, with six falsifiable hypotheses. Easier to falsify than to market, which is the point.
Read the paper →A neural network that reasons over structure rather than over tokens, with a formalized representation per level of abstraction. Knowledge as gravity rather than as retrieved text.
Read the description →An intelligence layer above agents: a task arrives in natural language and a verifiable artifact comes back with computed confidence, instead of a plausible paragraph.
Read the description →One repository, fifty questions in natural English, the model and its settings held constant, answers graded blind. Full configuration published so the test can be run again.
Read the benchmark →The unit and the six hypotheses, in full — with the PDF on the page.
Read the paper →A controlled study on whether structured meaning teaches a small model to reason about a system better than source code does. At inference it beat even the ideal source baseline; as training material at 3B it did not transfer. We will publish that half too.
In preparationGitMir's intelligence technology is proprietary. We publish benchmark methodology and measurable outcomes without publishing the internal representation or interpretation system — which is why every result here can be checked without us.
We collaborate with teams researching reasoning structures, program understanding and the economics of intelligence.