Same repository. Same questions. Same model. Each system received the context it normally provides. Answers were scored blind against verified repository evidence.
SAME AI. DIFFERENT INTELLIGENCE.
50 questions written as real engineering requests, not code-search queries.
| System | Answer quality | Context used | Cost / question | Latency |
|---|---|---|---|---|
| GitMir | 92% | 8.4K tokens | $0.07 | 4.8s |
| Unblocked | 78% | 27.4K tokens | $0.19 | 9.6s |
| Augment Code | 74% | 31.6K tokens | $0.22 | 10.9s |
| Sourcegraph Cody | 69% | 38.1K tokens | $0.26 | 12.4s |
| Qodo | 64% | 35.8K tokens | $0.24 | 11.7s |
| Greptile | 66% | 42.7K tokens | $0.29 | 13.1s |
| Raw repository context | 58% | 61.3K tokens | $0.41 | 18.9s |
RELATIVE IMPROVEMENTS ARE CALCULATED AGAINST UNBLOCKED, THE HIGHEST-SCORING ALTERNATIVE IN THIS RUN.
It shows how much scored answer quality each system produced per unit of context in this benchmark. It is computed from the table above, not measured separately.
THE MULTIPLE IS INDEXED TO UNBLOCKED = 1.00×, THE HIGHEST-SCORING ALTERNATIVE IN THIS RUN. DERIVED FROM THE MEASURED RESULTS ABOVE; NOT INDEPENDENTLY MEASURED.
GitMir scored highest in all four evaluated categories in this run.
| Question type | GitMir | Unblocked | Augment | Sourcegraph | Lead vs best alternative |
|---|---|---|---|---|---|
| Product understanding | 94% | 75% | 69% | 63% | +19 pts |
| Change impact | 91% | 74% | 71% | 68% | +17 pts |
| Multi-step behavior | 93% | 73% | 66% | 61% | +20 pts |
| Failure reasoning | 88% | 76% | 72% | 65% | +12 pts |
HEADLINE QUALITY IS CALCULATED OVER ALL 50 INDIVIDUAL QUESTIONS. CATEGORY SCORES ARE COMPUTED WITHIN EACH CATEGORY AND ARE NOT AVERAGED TO PRODUCE THE HEADLINE NUMBER.
Nobody assigns work by saying "find the definition of SubscriptionService". They say what they want to happen. Every question is written in the form a developer, product manager or operator could actually ask it.
What happens after a customer upgrades from the free plan to a paid plan?
What parts of the product could be affected if we allow customers to pause a subscription?
Add loyalty rewards to checkout. What needs to change?
What could break if refunds are allowed after an order has already shipped?
Why might a customer be charged successfully but still not receive access to the paid features?
What happens when a user deletes their organization?
What needs to change if we introduce a second approval step for large payments?
What customer-visible behavior depends on the current authentication flow?
A customer changed their email address. What downstream behavior could be affected?
What needs to happen across the product when a subscription expires?
10 of 50 shown. Full question set included in the published configuration.
We publish the repository, commit, model, settings, question set, scoring criteria and evaluation procedure required to reproduce the benchmark. We do not publish how GitMir produces the intelligence supplied to the model. That is proprietary technology.
Compare GitMir against the AI workflow you already use.
SAME MODEL. SAME TASK. MEASURE THE DIFFERENCE.