A bench for generalisation, run under a fixed budget.
Every network here is given the same storage and the same compute, then asked to learn a task it has never seen. A win that comes from a larger model or a longer run says nothing about the architecture, so those two are fixed before training starts.
One of the networks is mine, and it is the reason this bench exists. It is hypervector based: the learned state is dense real vector fields rather than layers of weights, and depth is resolution the task can use instead of work it must pay for. It is not published, nothing has been evaluated against it before this, and it performs surprisingly well. More on the network
The bench is not only a ranking. Each network has something it is good at, and the same runs that produce the ranking are used to show it: where a network gains, on which task, and at what point in the budget. The strengths and the advantages come out of the same empirical results, not from a separate argument.
The tasks split two ways. Self generalisation is games learned from self play, where the only signal is who won and it arrives once per game. Supervised is tasks with a label on every example, where a signal arrives at every step. Architectures do not fail evenly across that split, which is why both directories exist.
Size is counted in stored bits. A trained float weighs 32 bits, a survival mask weighs one bit per possibility it keeps, and state that can be regenerated from a seed is free because nothing was learned into it.
Compute is counted in scalar operations, from the first pass over the data to the last. Wall clock and the device that produced it sit beside that count, never in place of it.
Where a result is thin it says so. One seed is one seed, and a curve still rising at the end of its budget has not finished.
The first results are on the go page: Go, 5x5, from self play