Multi-agent reinforcement learning · In progress · Still training
TagTraining still in progress
Training is ongoing. Two learned agents play timed rounds of tag in a physical arena: one hides behind cover it can move, the other has to find it and touch it before the clock runs out. Set the round length, rearrange the arena, watch what each agent sees and remembers, and compare with the initial policies.
The best eligible evaluated policy pair plays a timed tag round in a shelter and doorway. Each agent receives its own sensors, its recurrent memory of the last sighting and the round clock; walls block its line of sight.The same frozen policies in connected rooms, with boxes, a plank and a ramp. Changing the layout tests their behavior without retraining; a tag ends the round and the next one starts on its own.
Source and documentation
The repository includes the implementation, reproducible examples, and tests.