Research
Papers
Published work, preprints, and what is currently open in front of me.
Research
Published work, and what is currently open in front of me.
The standard way to make a language model more reliable is to sample several reasoning traces and take the most common answer. On hard science questions, that makes it worse more often than it helps. Across 198 GPQA Diamond problems, majority voting reduced accuracy on 56.6% of problems for one model and 65.7% for another. The obvious response is to detect the bad cases and route around them, so I tested three cheap confidence signals that could do that. All three fail, each for a different reason, and one of those failures I still cannot explain. Every hypothesis was registered and git-tagged before the confirmatory data was touched. Version two revised the mechanism claim in my own accepted paper rather than defending it.
The follow-up
In progress
The unexplained failure above is what I am working on now. The question is whether agreement between samples measures evidence at all, or just measures how strongly the model already believed the answer before it reasoned. To test it I needed reasoning-model data that the published study could not afford to buy, so I generated it myself: 12,672 samples across 198 problems, 82.7 million tokens, twenty-three hours on a consumer laptop GPU. The hypotheses and kill conditions are written down and tagged before anything is run, so the result is whatever it is. No conclusion yet, and possibly no paper. That is the arrangement.
Upcoming Next
Divergence
An in-progress study of how reasoning models diverge from themselves: one prompt, one set of weights, sampled traces that reach incompatible conclusions. It follows a published result of mine on self-consistency failure modes and moves from measuring disagreement to characterising which internal states predict it. Targeting ICLR 2027. The harness and the dataset land before the paper does.
- ORCID
- 0009-0004-0034-824X
- Scholar
- Profile
- Entries
- 2