GitMir Science

Measuring what makes
AI more intelligent.

We study how information representation affects reasoning quality, context requirements, cost and reliability.

Areas of study

Six questions
we are working on.

Information Density

How much usable meaning a representation carries per unit of size — and how that changes what a model can do with it.

Reasoning Density

How much correct reasoning a model performs per unit of context it is given.

Semantic Distance

The distance between human intent and the information a machine is given. The larger it is, the more reasoning is spent reconstructing what was already meant.

Context Efficiency

How much of the context supplied to a model actually contributes to the answer.

Model Performance

How the same model's accuracy, cost and latency move when only the intelligence it receives changes.

Agent Economics

What one completed customer outcome costs across a production agent workload, rather than what one call costs.

Measurement

We do not only build intelligence.
We measure intelligence differently.

The industry measures models. We measure what a model is given. These are the units we report in, defined tightly enough to be computed the same way twice.

Context Information Density

Useful units of meaning per KB. How much a representation says in the space it takes.

Reasoning Efficiency

Correct task outcomes per unit of context. What a token of context is actually worth.

Semantic Distance

Distance between human intent and machine-readable information. How far a system has to travel before it can start reasoning.

11.0quality points per 1K tokens — 3.8× the highest-scoring alternative
92%answers correct, against 78%
8.4Ktokens per question, against 27.4K

FROM BENCHMARK #001. EVERY FIGURE WE PUBLISH CARRIES ITS METHOD, SCOPE AND LIMITATIONS.

Published

What we have measured so far.

Benchmark #001 · Published

Answer quality against context spent

One repository, fifty questions in natural English, the model and its settings held constant, answers graded blind. Full configuration published so the test can be run again.

Read the benchmark →
Paper · Published

Agentic Systems Load

The unit and the six hypotheses, in full — with the PDF on the page.

Read the paper →
Study · Write-up in preparation

Structured meaning at inference, and its limits in training

A controlled study on whether structured meaning teaches a small model to reason about a system better than source code does. At inference it beat even the ideal source baseline; as training material at 3B it did not transfer. We will publish that half too.

In preparation
Posture

We publish the results.
Not the recipe.

GitMir's intelligence technology is proprietary. We publish benchmark methodology and measurable outcomes without publishing the internal representation or interpretation system — which is why every result here can be checked without us.

Working on the same questions?

We collaborate with teams researching reasoning structures, program understanding and the economics of intelligence.