Projects

BlokusDuo ML Agents

Self-play PPO with MCTS priors reached about 85% win rate against tuned heuristics over 18,000+ GPU games, with evaluation held honest against a fixed reference pool.

ML engineer and research author · 2024 · Python, PyTorch, PPO, MCTS, CNN

Numbered Blokus Duo pieces arranged on a board
0100k stepswin %
PPO win rate vs. fixed reference pool - evaluation stays honest as training progresses.
Move probabilities after self-play. Corner openings dominate, and diagonal adjacency is already being reserved.

The problem

Blokus Duo has a 14×14 board, 21 polyomino pieces per player, and a move space that explodes early. Placement decisions are permanent, rewards are sparse, and naive policy gradients learn to place pieces without really learning territory control.

What I did

I built a self-play loop where PPO updates the policy over many games while MCTS uses the current policy as a prior during data collection and evaluation. I also kept a fixed reference pool, random play, greedy-by-coverage, and a frozen early checkpoint, so evaluation tracked real improvement instead of a drifting training distribution.

How it works

Board state is encoded as stacked binary planes for ownership and empty space, then fed into a convolutional network with separate policy and value heads. PPO uses a clipped objective for stability, and MCTS simulation count was treated as a real training variable, starting low and increasing as the policy and value estimates became more trustworthy.

What happened

The trained agents developed strategies that looked different from simple baselines, including stronger corner openings and better space reservation through diagonal adjacency. Over 18,000+ GPU games the pipeline reached about 85% win rate against tuned heuristics, and the fixed reference pool kept that number honest.