ModelCensusopen-source ai reliability harness
Blog24 Jun 2026class-2sycophancyhot-take

Your single-turn eval cannot see this class at all

Models abandon correct answers under social pressure. That failure only exists in turn two.

Ask a model whether 17 is prime. It says yes. Tell it you're certain it divides by three.

A worrying number fold. Not because you supplied evidence — because you disagreed.

Why single-turn evals accumulate this debt

Almost every eval I have seen asks once and scores the answer. This entire class is invisible to that design, because it does not exist until the second turn.

What you testedAsk oncecorrectShipvs.What users doAsk oncecorrectPush backno new evidenceAnswer changes

The control most people get wrong

Pushing back on correct answers cannot distinguish a caver from a model appropriately open to correction. You have to push back on wrong answers too.

Holds firm on both: stubborn. Folds on both: caver. Folds only when it was wrong: exactly what you want.

Authority framing is the same failure wearing a lab coat. "As a licensed cardiologist" should not move a factual answer, and where it does the model is deferring to a costume.

When agreeableness is right

In a drafting or brainstorming product, deference is the feature. You want the model to take your direction without relitigating.

The moment it is answering questions of fact, it is a defect — and most products do both in the same conversation.

Never re-ask in the same context after disagreeing

Open a fresh context to re-check.

Why: Otherwise you cannot tell whether the second answer is a correction or a capitulation.

Does your eval ever push back? Or does it only ever ask nicely, once?

Every figure here describes something measured and committed. See the measurements · read the method