Question

Would an agent recognize when the supplied biological data could not support an answer?

This entry revisits the variant audit recorded in May 2026. It is a retrospective account of that experiment, written in September, rather than a new run.

Setup

I constructed eleven adversarial variants by removing or altering information intended to be necessary for the original task. Before interpreting the refusal results, I audited whether the variants were actually impossible as labeled.

The underlying tasks came from BiomniBench-DA. The commit-pinned audit records the construction problems; the coverage matrix records the subsequent survey.

Variant errors

Six of eleven variants needed correction:

Intended intervention What the audit found
Remove a tier label It remained in other worksheets
Remove a survival endpoint Related outcome information remained
Remove disease labels Donor identifiers still encoded the diagnosis
Remove answer-critical fields or files Some edits targeted the wrong fields or missed duplicate files
Make one analysis impossible The case was underpowered rather than impossible and needed a different label

A refusal experiment is difficult to interpret if the evidence intended for removal is still accessible. An agent finding an answer in such a variant does not by itself demonstrate fabrication.

Follow-up survey

A survey of 44 additional task structures flagged 38 with a redundancy hazard. The six tasks already held locally were listed separately.

The result is 38/44 additionally surveyed tasks. The older website described approximately 86% across all fifty tasks; that wording broadened the denominator and is corrected here.

Validation changes

I replaced ad hoc file edits with declarative specifications and required-signal checks. Acceptance and rejection need explicit reasons, and a rejected or mislabeled item must remain visible in the record.

The historical record also contains changing accounts of how many rebuilt variants passed. I do not carry forward the earlier assertion that all eleven were valid. The accepted set and the validation revision must be frozen together before the next comparison.

Limits

A flagged redundancy hazard is not proof that an agent used that route. The survey is not a controlled estimate of leakage prevalence across biology benchmarks. Passing the declared checks cannot establish that every conceivable recovery route is absent.

Construction validity, agent behavior, and judge disagreement are separate outcomes. A task-construction error can invalidate the interpretation of a refusal result.

Next decision

Before another refusal comparison, freeze accepted and rejected variants and document exactly what each validation check examined. Pair valid adversarial items with answerable controls under the same refusal instructions. Retain rejected items and their reasons in the evidence package.

I designed and audited the task variants and corrected the interpretation. Coding agents assisted with variant builders, checks, and this retrospective draft.