GoSearch

Go, 5x5, learned from self play.

The rules are the full ones on a small board: captures, no suicide, positional superko, area scoring, komi 7.5, and the game ends after two passes or 60 moves. No human games and no handcrafted evaluation go in. The only signal is who won.

Training is alphazero style. A network with a policy head and a value head guides a PUCT search, the search visits are the policy target, the result of the game is the value target, and every position is used under all 8 board symmetries. The budget, what each run consumed of it, and the time it took are in the first table.

Three networks ran, matched on the part that does the work. The hypervector network holds 15,131 parameters, 8,192 of them in its core. The mlp holds 14,890, 7,951 in its core. The third is the two heads alone with no core at all, 6,939 parameters, and it is there as a floor. Every run and every match was made on the same machine, Apple silicon with 10 threads.

Against a fixed outside opponent, classical MCTS with 1,000 random playouts, the three score 0.80, 0.78 and 0.37 with 200 simulations of their own. Against 4,000 playouts they score 0.65, 0.495 and 0.075.

With the search taken away the gap widens. Playing its policy alone, one forward pass and no lookahead, the hypervector network scores 0.59 against 1,000 playouts. The mlp scores 0.325 and the heads alone 0.20. It is the only one of the three whose raw policy beats that opponent.

The networks were then played against each other, 200 games a match with colours alternated, in the second table. The last three rows are against the exclusion network from the gahp project, which keeps no weights at all, only a record of the moves it has ruled out in 1024 pattern tables. That model holds 3.33 million survival bits against 15,131 parameters here, and was trained on 10,000 self play games.

Under the exclusion network's own rules, komi 0.5 and a 75 move cap, those three matches give 0.805, 0.785 and 0.490 instead. The advantage holds on settings the network was not trained for.

The final network was also played against its own earlier checkpoints, 200 games each with search on both sides: 0.96 against iteration 10, 1.00 against 30 and 50, 0.68 against 80, 0.765 against 100, and 0.50 against 130. It never loses to an older version of itself.

NetworkBudgetConsumedSelf-playTrainingWall clock
1Hypervector30,000 games x 200 sims1.212 M positions876 s79 s1,269 s
2MLP30,000 games x 200 sims1.145 M positions572 s71 s957 s
3Heads only30,000 games x 200 sims1.191 M positions495 s45 s829 s
MatchModeGamesScoreRecord
1Hypervector vs MLPraw policy2000.685137-63
2Hypervector vs MLP200 sims each2000.810162-38
3Hypervector vs MLP200 sims, 8 open plies2000.625125-75
4Hypervector vs MLPraw vs 200 sims2000.20541-159
5Hypervector vs Heads onlyraw policy2001.000200-0
6Hypervector vs Heads only200 sims each2001.000200-0
7MLP vs Heads only200 sims each2001.000200-0
8Hypervector vs Exclusionraw policy2000.860172-28
9Hypervector vs Exclusion200 sims each2000.955191-9
10Hypervector vs Exclusionraw vs 200 sims2000.735147-53

What the results say about the architectures.

The hypervector core is what the policy is made of. At a matched core size its raw policy beats the mlp 0.685, and it is the only policy here that beats 1,000 playouts of search on its own. The heads alone, the same encoder and the same two heads with no core between them, score 0.37 with search and 0.20 without it, so nearly all of that play comes from the core rather than from the search around it.

Search hides the difference. With 200 simulations on both sides the hypervector network wins 0.81 against the mlp, yet against the rollout opponent the two score 0.80 and 0.78, nearly level. Lookahead repairs a weaker policy, and the ranking only separates cleanly once the lookahead is taken away.

That cuts both ways. The hypervector network scores 0.205 when it gives up its own search and the mlp keeps 200 simulations, so the policy advantage is worth less than the search it gave up.

The exclusion network is the opposite shape. It stores 220 times the bits and had a search of its own, and it still loses 0.86 on raw policy, because bans accumulate slowly when the only signal is who won a game. An exclusion-only learner was also given this same alphazero setup, with policy and value heads built from bans, and at tau 0.5 its policy collapsed inside one iteration, nearly every move banned in every pattern. That run was stopped rather than tuned, and is not counted here.

What the runs do not settle. One seed per network, and with 200 games a match one standard error is about 3.5 points. Komi 7.5 favours black heavily at this level, so between two close networks colour decides most games, which is why the checkpoint ladder reads 0.50 at the last rung. The learning rates and the search constants are the defaults, not tuned per network.

More on the architectures: Reverire and GAHP