BK

Building an AlphaZero-style chess engine in Python

No opening book and no human games. It plays itself, and it learns from the result.

By Bhargavaram Krishnapur4 min readMachine learning

Painting: Albert Bierstadt, Mount Starr King, Yosemite, 1866. Cleveland Museum of Art, CC0.

AlphaZero is famous because it learned chess with no human knowledge beyond the rules. It played itself on very large hardware, and it became stronger than every engine before it. I wanted to know how much of that idea still works on one ordinary computer, in code that a person can read in an afternoon.

AZ-Lite is the result. It is a compact engine in Python. It combines Monte Carlo tree search with a small neural network, and it learns from self-play. It is built to be readable, reproducible and runnable locally. It is not built to beat Stockfish.

The pieces

The engine has four parts, and each one is small.

  • The board encoder. A position becomes a 12 × 8 × 8 tensor: one 8 × 8 plane for each of the six piece types in each of the two colours. A small three-layer convolutional network turns that tensor into a compact embedding of the position.
  • The policy head. Chess has a very large set of possible moves, and most of them are illegal in any given position. So the engine does not predict over a fixed list. It learns embeddings for the from square, the to square and the promotion piece, combines them with the position embedding, and scores only the legal moves, on the fly.
  • The value head. One number between −1 and 1 that predicts the result of the game from this position.
  • The search. Monte Carlo tree search with the PUCT rule. The policy tells the search which moves to look at first. The value tells it how good a position is without playing the game to the end.
Self-playPUCT tree searchDirichlet noise at the rootReplay data(state, π, z) examplesone line of JSON eachTrainingvalue: mean squared errorpolicy: cross-entropyNew checkpointa better networkguides the next searchone loop
The AZ-Lite training loop. Each turn of the loop gives the search a better guide.

The loop

Training is a closed loop with three steps.

Self-play. The engine plays games against itself. At each move it runs a tree search and records the position, the visit counts of the search (called π), and later the final result of the game (called z). Dirichlet noise is added at the root of the search, so that the engine also tries moves it would otherwise ignore. Each example is saved as one line of JSON.

Training. The network samples from the saved examples. It learns to make its value match the real result and its policy match the search. The loss is the mean squared error on the value plus the cross-entropy on the policy.

A new checkpoint. The next round of self-play uses the improved network. The search gets better guidance, so the games get better, so the training data gets better.

That is the whole idea of AlphaZero, and it fits in one file. The command line has three modes: selfplay, train and play. In play mode you type moves such as e4 or Nf3, and the engine answers.

Mirror Case with a Couple Playing Chess, 1325–50. Cleveland Museum of Art, CC0.

What small hardware changes

The real AlphaZero used thousands of processors. On one CPU, the number of self-play games is the limit. With a small dataset, the loss falls slowly and the network learns slowly. That is the honest cost of running everything locally.

It also changes what the project is for. AZ-Lite is not a strong engine. It is a working, readable version of the method, where each part can be changed and tested on its own.

What comes next

The roadmap is public in the repository. The next steps are a configuration file for the hyperparameters, more tests, and profiling of the self-play and training loops. After that come support for the UCI protocol, so that the engine works in standard chess programs, a larger network, and a pipeline that measures the engine's strength as an Elo rating over time.

The work this is about

Bhargavaram Krishnapur
Bhargavaram Krishnapur

Computer science student at Vijaybhoomi University and founder of The Pulse Engine. I build local-first, open-source tools. See the portfolio or GitHub.