2 articles on benchmarks.

A practical protocol for testing AI code tools yourself: designing the question set, holding the model constant, blind scoring, and the four things worth measuring.

Vendor architecture slides are unfalsifiable, and treating them as evidence is the most common mistake in AI procurement. What counts as evidence, and six questions for a vendor call.