Pictura: Perspective-View Self-Play
at Scale for Driving

Driving policies trained by self-play directly from rendered camera views, with no privileged observation of the surroundings.

Yuan Yin1, Elias Ramzi1, Marc Lafon1, Valentin Charraut2, Victor Bares2, Yihong Xu1, Éloi Zablocki1, Alexandre Boulch1, Thibault Buhet2, Andrei Bursuc1, Matthieu Cord1,3

One frame the policy learns from. The labels mark the scene elements Pictura supports beyond moving vehicles: traffic lights, pedestrians, parked cars and walls are all simulation primitives drawn by its CUDA rasterizer, inside the training loop.

The scale of one self-play training

50 Bagent steps
200 Bcamera frames renderedfour-camera rig
35 M kmdriven≈ 91 trips to the Moon
1750 yearsof driving at 20 K km/yearthe same distance, in human terms
under 13 h on 32 H100sat ~1.3 M agent-steps/s on average

How it works

Every agent acts, is rendered for its neighbors, and learns from what it sees, all within one closed loop. Perception and control are learned jointly.

Pictura architecture: a vectorized simulator sends state to a custom PV rasterizer on GPU, which renders each agent's four camera views; those observations go to the Alberti policy, whose actions step the simulator, while observation, action and reward feed the PPO update.

Pictura couples a vectorized simulator with a custom GPU rasterizer that renders the simulator state into each agent's egocentric perspective view. The observation feeds Alberti, the policy, and, together with the action and the reward, forms the rollout that PPO turns into a policy update.

How fast the renderer is

Perspective self-play is practical only if rendering is cheap. Pictura's CUDA rasterizer runs ahead of state-of-the-art perspective renderers, and it gets faster on newer hardware optimized for deep-learning training.

A new state of the art that scales with hardware: 1.4–4.1× over the best prior setting, and 1.5× again from A100L to H100.

Render-only agent-steps per second against render resolution, on one A100L. Pictura's CUDA rasterizer runs from 333K at 64 pixels down to 64K at 512, above the prior batched Madrona rasterizer and ray tracer and above Pictura reimplemented on Madrona, while HUGSIM and RAP sit near zero. In the right panel, moving from A100L to H100 lifts the CUDA rasterizer to 501K-102K while the Madrona reimplementation drops.
Render-only agent-steps/s against render resolution. Left: every system on the same A100L, at the square wall-free setting the starred numbers are reported under. Right: our CUDA rasterizer and its Madrona reimplementation from A100L to H100, where that same rasterizer on Madrona instead slows to 0.2–0.7×.

No longer the bottleneck: rendering takes about 10% of a training step, against more than half on Madrona.

Wall-clock of one training iteration by phase, in seconds and as a share of the iteration, for render resolutions from 32 to 384 pixels. Rendering is a thin slice for the CUDA rasterizer and over half for Madrona, which runs out of memory at 384 pixels.
Per-phase wall-clock of one training iteration on a single H100, in seconds and as a share of the iteration. The Madrona path makes the whole iteration 2–3× slower and runs out of memory altogether at 384×216.

Watching the policy drive

Each clip is 25 seconds of self-play by Alberti 50 B, showing exactly what the policy receives: the four cameras of its rig, rendered at 96×54 pixels by Pictura inside the training loop. Within a family every face carries its own hue, so an object's heading reads from its shading. Click any clip for a higher-resolution render, for viewing only.

vehicles pedestrians cyclists buildings road edges road lines crosswalks lane centers

HD
Town10HD · downtown junction, a pedestrian and cyclists
HD
Town02 · a pedestrian and cyclist ahead
HD
Town07 · junction, pedestrian and cyclist among the cars
HD
Town02 · following traffic, vehicles ahead and behind
HD
Town01 · junction with cyclists and a pedestrian
HD
Town03 · cyclists among the vehicles

Traffic density

These clips show Alberti driving in Town10HD across light, medium and dense traffic.

HD
60 agents · light
HD
120 agents · medium, the training density
HD
150 agents · dense

Grounding in what the cameras see

Two probes ask what Alberti and a privileged vectorized policy (same training recipe, but reading exact state) actually base their decisions on.

Counterfactual probes: Alberti's response falls as occlusion rises and reaches zero once the road user is hidden, while the privileged policy stays sensitive to agents no camera could see.

Aggregate Magnitude of the change in value and in braking probability caused by deleting one nearby road user, against how occluded that user is. Alberti's response falls steadily to zero at full occlusion, while the privileged vectorized policy stays flat and high.
Example A pedestrian swept along a path behind a vehicle: a top-down view of the path colored by visibility, and the signed responses read along it. Alberti reacts as the pedestrian comes into view and returns to zero when it is hidden, while the vectorized policy responds throughout.
Change in value ΔV and in braking probability ΔP(brake) when one nearby road user is removed from the observation and the rest of the scene is left untouched. Left: over 88 K removals in 16 scenes, against how occluded the removed user is. Right: one pedestrian swept along a path through a frozen scene, the path colored by visibility (green to red, fully visible to hidden).

Blind corners: Alberti approaches a corner hiding an oncoming car more slowly, so it is further away when the car appears — caution a privileged expert has no reason to learn.

Aggregate Mean ego speed after takeover. With an oncoming car hidden behind a corner, Alberti slows earlier and stays slower than the privileged vectorized agent, which is faster until the car comes into view. At comparable corners with nothing hidden, the vectorized agent accelerates while Alberti hedges anyway.
Example One example takeover at a building-occluded junction, with the moment the hidden car enters each agent's input and the moment it comes into line of sight marked on the speed curves.
Ego speed after each policy takes over the same recorded state, every other agent replaying its logged trajectory. Left: mean speed with an occluded oncoming vehicle and at comparable empty blind corners; the gray band marks the main reveal times. Right: one takeover with cross traffic occluded by a building — ×: the hidden car first enters that agent's input; ○: it first comes into line of sight.