Keep It Together
Six short acts. Jump around, or replay the walkthrough.
Hold the keep with differently-trained AIs.
Budget 0 left
Your best hold, by partner
Finish a game and it is added here. Only your best hold with each partner is kept, and nothing is saved when you leave.
A catapult in the partner's gold: a picture of the pieces it puts up, not of the engineer.
This partner trained only with a copy of itself.
Your own row needs a fortress to play with. Build one in Your turn and it appears here.
Takes the game score out of the training and leaves only the part that asks for six plans you can tell apart.
Two engineers hold one keep. Each has its own budget, neither can finish a line alone, and they cannot talk to each other. Everything you see is played live in your browser by small networks I trained for this game. One of them trained with a copy of itself, and it is very good, right until it meets an engineer that learned a different plan. The other trained against a partner that kept changing its plan, so it spends the first seconds of a game working out what you are building before it commits. What that bought, and what it did not, is measured, and the numbers are in the fine print. This is an illustration of Any-Play, a paper I wrote with Ross Allen in 2022, not the experiment in it.
A caveat: this is a personal side project and I have not tested it exhaustively. Every number here comes from games your browser plays while you watch, so yours will not match mine exactly.
Original game, art and code. The idea comes from Any-Play (Lucas and Allen, AAMAS 2022).
How this page works
Your browser asks for reduced motion, so no wave on this page animates. Watch a wave, Reading the plan and each wave you send in Your turn are shown finished, and Step advances one a move at a time. In Strangers every square is a game played to the end. Six plans and Fine print have nothing that moves.
The game
Two engineers hold one keep. Each spends its own budget, they build at the same time, and they cannot talk to each other. Walls cost 2 and catapults cost 10. Three waves of walkers come in through the gates. Whatever is left of the keep at the end, out of 100, is the score.
The controls
- In Your turn, tap or click a tile to place the piece you have selected. A ghost under the pointer shows whether that tile will take it.
- Wall and Catapult pick the piece. Undo takes back your last placement. Done building ends your build and starts the wave.
- During a wave, tap a walker that has stopped to call the shot: your own machines in reach will prefer it.
- Hover a machine, or tap one on a phone, to open the close up.
- Keyboard: arrow keys move the cursor, Enter places, Backspace removes, W picks the wall, M picks the catapult, S mutes the sound, Escape backs out.
The six acts
- Watch a wave. One pair that trained together builds and holds, with nothing asked of you.
- Six plans. One network run six times, once for each hidden number it was trained on, played as six games side by side.
- Strangers. Every engineer paired with every other. Down the diagonal they trained together; everywhere else they are strangers.
- Reading the plan. One pairing played large, with the reader's own guess about its partner's plan drawn live beside the wall.
- Your turn. You take one seat and an AI takes the other. The partner picker decides which one.
- Fine print. What is measured here, what comes from the paper, and what is mine.
Fine print
How honest is this?
This is a game I made up for this page. It is not the card game the paper ran its experiments on, and it is not those experiments. The paper works in a two-player co-operative card game where you can see everyone's hand except your own. A wall is easier to look at than a card game, and the thing being shown is the same thing, but the wall is an analogy and the numbers on this page are this game's numbers.
The engineers are small networks trained on my machine and shipped here as one file. Nothing is sent anywhere. Every game you see is played in your browser while you watch, which is also why your numbers will not match the ones I saw.
The six-plan engineer is one network, run with a different hidden number each game. It is not six engineers, and it is not what ships as a teammate. In the paper it is scaffolding: its whole job is to be a varied training partner, and then it is thrown away. The engineer you play with in Your turn is the one that trained against it.
Nothing here is sneaky. No engineer decided to keep a secret. When two networks train together for long enough, they settle into a shared way of doing things because nothing ever pushed them to do otherwise, and neither of them has any idea that the way is one of many. The word for it in the literature is a convention, and it is an accident of training rather than a plan.
More plans is not automatically better. In the small game the paper studies closely, exactly four secret numbers produced agents that paired perfectly with every other agent trained the same way, and three, five and six all left gaps. The right number depends on the game, and six is the number that fits this map.
The map never changes, but each game does. Which gates open, how hard they push, where the rubble lands, what is left of the old walls, and how much budget there is all move from game to game. Without that, an engineer could tell its partner's plan from a single memorised tile, and it would have learned nothing about reading plans.
Two other methods come up if you read around this. Other-Play and Off-Belief Learning go the other way: instead of making many plans, they remove the arbitrary part so that every training run lands in the same place. Within their own family they do that very well, and Off-Belief Learning was the strongest partner anyone had for a human-like teammate at the time. The gap this page is about opens when the stranger came out of a different algorithm altogether.
Measuring that gap, pairing agents across algorithms rather than across seeds, is what the paper proposed. Four years on it is still not something the field reports as standard, so take it as a lens I argued for rather than as settled practice.
The grid of strangers is one game per square, which is enough to see the shape and not enough to be a measurement.
Any one game is one game, and one game is noisy.
From the paper: the card game, two players, out of 25
| Paired with the partner it trained with | Paired with an agent from a different algorithm | |
|---|---|---|
| Off-Belief Learning | 24.2 | 5.2 |
| Any-Play, on SAD+AUX | 22.5 | 14.2 |
The Any-Play agents score lower with their own kind and nearly three times higher with strangers. That trade is the whole point, and it is why "how good is it" is the wrong single question to ask about a teammate.
Against partners that were never built for teaming with strangers at all, the same two scored 1.0 and 7.4.
These are points from the card game, in Table 1 of the paper. Nothing else on this page is measured in them, and the two scales should not be put on one axis.
Being good with your twin, against being good with strangers
| With your own twin, against strangers from other algorithms | -0.23 |
|---|---|
| With your own twin, against partners never built for strangers | -0.55 |
| With one kind of stranger, against the other | +0.89 |
Practising with your twin is not the same as getting good with strangers. Across the pool in the paper, the two ran in opposite directions: the more an agent scored with its own kind, the less it tended to score with everyone else. Pearson correlation -0.55.
A separate study is the reason I care about this. Siu and colleagues, in 2021, sat people down to play with two AI teammates that scored the same. On all eight questions they asked afterwards, people preferred the teammate whose moves they could follow. Scoring well and being a good teammate turned out to be two different things.
That study tested a different technique from this one, so it is not a scoreboard between them. I read it as the reason the whole line of work exists.
Common questions
Is this the card game from the paper?
No. What the paper measured, it measured in a two-player co-operative card game. This is a wall-defence game I wrote for this page, because a wall with a hole in it is easier to look at than a discarded five. The mechanism is the same one, the analogy is mine, and every number on this page other than the ones on the labelled 25-point card came out of this game.
Why does the stranger fail? It looks like a good engineer.
It is a good engineer. So is the one next to it. Each of them learned a plan that works when a partner builds the other half, and nothing during their training ever told them which plan the rest of the world uses, because there was no rest of the world: each of them trained with a copy of itself. Put two of them together and you get two competent halves of two different plans, with a gap where they meet.
Are those six separate AIs?
No, it is one network run six times. Before each game it is handed a number from one to six that its partner cannot see, and it learned to build a different way for each number. The reason it bothered is that it was trained to make its plan recognisable to its partner, on top of the usual reward for holding the wall.
Would more plans be better?
Not automatically. In the small game the paper looks at closely, four secret numbers produced agents that paired perfectly with every other agent trained the same way, and three, five and six all left gaps. The right number depends on the game. Six is the number that fits this map, and it was chosen by testing rather than by reasoning that more must be better.
Did the AI name the plans?
No, and neither did I. It was handed a number and no words at all. The names are written by the same code that scored the six: it reads where each style actually spends its credits across a fixed set of test games and names the patch of the board that style favours. If you think one of them is misnamed, the layouts are the thing to trust rather than the label.
Credits
The idea comes from Any-Play, a paper I wrote with Ross Allen for AAMAS 2022. You can read the paper, the code, a short video, and the article MIT wrote about it.
The reward for being recognisable is adapted from DIAYN (Eysenbach and colleagues, 2019), which learns a set of distinguishable behaviours with no score at all. Keeping the score is the change Any-Play makes, and it is the difference between six ways of winning and six ways of being unusual.
Training a partner against a frozen cast of others comes from Fictitious Co-Play (Strouse and colleagues, 2021). Its cast comes from random seeds and saved checkpoints rather than from any objective that asks for variety, which is the gap Any-Play sets out to close.
Other-Play and Off-Belief Learning (Hu, Lerer, Foerster and colleagues) take the opposite route, removing the arbitrary part so independently trained agents land in the same place. They do that well with their own kind. TrajeDi (Lupu and colleagues, 2021) pushes a population of policies apart with an explicit diversity term. The benchmark that made all of this measurable is the card game challenge of Bard and colleagues. The human study is Siu and colleagues, 2021. The failure where a diversity reward is satisfied by playing badly in a distinctive way is named in ADVERSITY (Cui and colleagues, 2023).
The game, the art, the training code and everything running on this page are original.