ECCV Workshop EmergingAD 2026
Self-play in simulation produces robust driving policies at scale, but demonstrations of such behavior rely on privileged vectorized observations such as exact poses and velocities, even for occluded agents. We establish perspective-view self-play as a practical training regime with Pictura, a GPU-accelerated multi-agent driving simulator that renders each agent's egocentric view at every step, sustaining up to 500K agent-steps/s (2M images/s) on a single H100. Using Pictura, we train Alberti by self-play with plain PPO: the first large-scale driving self-play policy trained directly from perspective images, without privileged observations. Training spans 50B agent steps for ~35M km of driving, approaching the driving performance of its privileged vectorized counterpart and transferring zero-shot to Waymo Open Motion Dataset layouts.