Go, 5x5, learned from self play.
The rules are the full ones on a small board: captures, no suicide, positional superko, area scoring, komi 7.5, and the game ends after two passes or 60 moves. No human games and no handcrafted evaluation go in. The only signal is who won.
Training is alphazero style. A network with a policy head and a value head guides a PUCT search, the search visits are the policy target, the result of the game is the value target, and every position is used under all 8 board symmetries. The budget, what each run consumed of it, and the time it took are in the first table.
Three networks ran, matched on the part that does the work. The hypervector network holds 15,131 parameters, 8,192 of them in its core. The mlp holds 14,890, 7,951 in its core. The third is the two heads alone with no core at all, 6,939 parameters, and it is there as a floor. Every run and every match was made on the same machine, Apple silicon with 10 threads.
Against a fixed outside opponent, classical MCTS with 1,000 random playouts, the three score 0.80, 0.78 and 0.37 with 200 simulations of their own. Against 4,000 playouts they score 0.65, 0.495 and 0.075.
With the search taken away the gap widens. Playing its policy alone, one forward pass and no lookahead, the hypervector network scores 0.59 against 1,000 playouts. The mlp scores 0.325 and the heads alone 0.20. It is the only one of the three whose raw policy beats that opponent.
The networks were then played against each other, 200 games a match with colours alternated, in the second table. The last three rows are against the exclusion network from the gahp project, which keeps no weights at all, only a record of the moves it has ruled out in 1024 pattern tables. That model holds 3.33 million survival bits against 15,131 parameters here, and was trained on 10,000 self play games.
Under the exclusion network's own rules, komi 0.5 and a 75 move cap, those three matches give 0.805, 0.785 and 0.490 instead. The advantage holds on settings the network was not trained for.
The final network was also played against its own earlier checkpoints, 200 games each with search on both sides: 0.96 against iteration 10, 1.00 against 30 and 50, 0.68 against 80, 0.765 against 100, and 0.50 against 130. It never loses to an older version of itself.