BlokusDuo ML Agents
Self-play PPO with MCTS priors reached about 85% win rate against tuned heuristics over 18,000+ GPU games, with evaluation held honest against a fixed reference pool.
The problem
Blokus Duo has a 14×14 board, 21 polyomino pieces per player, and a move space that explodes early. Placement decisions are permanent, rewards are sparse, and naive policy gradients learn to place pieces without really learning territory control.
What I did
I built a self-play loop where PPO updates the policy over many games while MCTS uses the current policy as a prior during data collection and evaluation. I also kept a fixed reference pool, random play, greedy-by-coverage, and a frozen early checkpoint, so evaluation tracked real improvement instead of a drifting training distribution.
How it works
Board state is encoded as stacked binary planes for ownership and empty space, then fed into a convolutional network with separate policy and value heads. PPO uses a clipped objective for stability, and MCTS simulation count was treated as a real training variable, starting low and increasing as the policy and value estimates became more trustworthy.
What happened
The trained agents developed strategies that looked different from simple baselines, including stronger corner openings and better space reservation through diagonal adjacency. Over 18,000+ GPU games the pipeline reached about 85% win rate against tuned heuristics, and the fixed reference pool kept that number honest.