Deep RL for Bazar Blot
Bazar Blot is the Armenian version of Belote, and the card game I grew up playing with my family. I'm training agents to play it with self-play reinforcement learning.
It's a hard problem to learn. Four players in two teams, 24 of the 32 cards hidden, a bidding round that sets the contract, and a partner you can only talk to through your bids and the cards you play. How much you should bid depends on how well you'll play, and how you should play depends on what was bid. The two have to be learned together.
What's built
- A rules engine for the Blot Star ruleset, with 330 tests, checked against a million random deals.
- A solver that finds the best possible play when all four hands are visible. I use it as ground truth.
- A training environment, plus baseline bots: random, a rule-based player, and one that samples possible hidden hands and solves each.
- An evaluation setup that replays the same deals with the seats swapped, so luck cancels out.
- A PPO self-play agent, and a web app where you can watch the bots or play against them.
- +67
- pts/deal vs heuristic
- +206
- pts/deal vs random
- 330
- tests, 1M-deal check
After about 28 minutes of training on one CPU core, the agent beats the rule-based bot by about 67 points a deal and random play by about 206. The confidence intervals stayed above zero at every check during training.
My first self-play run broke: all four seats learned to always pass, so no hand was ever played. I fixed it with two changes. A round where everyone passes now counts as a real zero-point result instead of being thrown away, and the opponent pool includes fixed baseline bots. I haven't yet checked whether its bidding is actually good, or only good enough to beat weak opponents.
Next: try different ways of training the bidder and the player together, then add search on top of the learned policy.
