Benchmark 001

Human-intent
software reasoning.

Same repository. Same questions. Same model. Each system received the context it normally provides. Answers were scored blind against verified repository evidence.

SAME AI. DIFFERENT INTELLIGENCE.

Head to head

Same model. Same task.
Different result.

50 questions written as real engineering requests, not code-search queries.

SystemAnswer qualityContext usedCost / questionLatency
GitMir92%8.4K tokens$0.074.8s
Unblocked78%27.4K tokens$0.199.6s
Augment Code74%31.6K tokens$0.2210.9s
Sourcegraph Cody69%38.1K tokens$0.2612.4s
Qodo64%35.8K tokens$0.2411.7s
Greptile66%42.7K tokens$0.2913.1s
Raw repository context58%61.3K tokens$0.4118.9s
+14 ptsanswer quality (18% higher)
69%context required
63%cost per question
50%latency

RELATIVE IMPROVEMENTS ARE CALCULATED AGAINST UNBLOCKED, THE HIGHEST-SCORING ALTERNATIVE IN THIS RUN.

Reasoning efficiency

A derived metric:
answer quality divided by context used.

It shows how much scored answer quality each system produced per unit of context in this benchmark. It is computed from the table above, not measured separately.

GitMir92% QUALITY ON 8.4K TOKENS
10.95QUALITY POINTS / 1K TOKENS
3.85×
Unblocked78% QUALITY ON 27.4K TOKENS
2.85QUALITY POINTS / 1K TOKENS
1.00×
Augment Code74% QUALITY ON 31.6K TOKENS
2.34QUALITY POINTS / 1K TOKENS
0.82×
Sourcegraph Cody69% QUALITY ON 38.1K TOKENS
1.81QUALITY POINTS / 1K TOKENS
0.64×
Qodo64% QUALITY ON 35.8K TOKENS
1.79QUALITY POINTS / 1K TOKENS
0.63×
Greptile66% QUALITY ON 42.7K TOKENS
1.55QUALITY POINTS / 1K TOKENS
0.54×
Raw repository context58% QUALITY ON 61.3K TOKENS
0.95QUALITY POINTS / 1K TOKENS
0.33×

THE MULTIPLE IS INDEXED TO UNBLOCKED = 1.00×, THE HIGHEST-SCORING ALTERNATIVE IN THIS RUN. DERIVED FROM THE MEASURED RESULTS ABOVE; NOT INDEPENDENTLY MEASURED.

3.8×more answer quality per 1K tokens than the highest-scoring alternative
By question type

Results by question type.

GitMir scored highest in all four evaluated categories in this run.

Question typeGitMirUnblockedAugmentSourcegraphLead vs best alternative
Product understanding94%75%69%63%+19 pts
Change impact91%74%71%68%+17 pts
Multi-step behavior93%73%66%61%+20 pts
Failure reasoning88%76%72%65%+12 pts

HEADLINE QUALITY IS CALCULATED OVER ALL 50 INDIVIDUAL QUESTIONS. CATEGORY SCORES ARE COMPUTED WITHIN EACH CATEGORY AND ARE NOT AVERAGED TO PRODUCE THE HEADLINE NUMBER.

Human-intent question set

Questions a person asks.
Not queries a search engine answers.

Nobody assigns work by saying "find the definition of SubscriptionService". They say what they want to happen. Every question is written in the form a developer, product manager or operator could actually ask it.

01

What happens after a customer upgrades from the free plan to a paid plan?

02

What parts of the product could be affected if we allow customers to pause a subscription?

03

Add loyalty rewards to checkout. What needs to change?

04

What could break if refunds are allowed after an order has already shipped?

05

Why might a customer be charged successfully but still not receive access to the paid features?

06

What happens when a user deletes their organization?

07

What needs to change if we introduce a second approval step for large payments?

08

What customer-visible behavior depends on the current authentication flow?

09

A customer changed their email address. What downstream behavior could be affected?

10

What needs to happen across the product when a subscription expires?

10 of 50 shown. Full question set included in the published configuration.

Methodology

The benchmark is reproducible.
GitMir is proprietary.

We publish the repository, commit, model, settings, question set, scoring criteria and evaluation procedure required to reproduce the benchmark. We do not publish how GitMir produces the intelligence supplied to the model. That is proprietary technology.

Run the same test
on your repository.

Compare GitMir against the AI workflow you already use.

SAME MODEL. SAME TASK. MEASURE THE DIFFERENCE.