Most agents that learn from video need to know what action produced each frame. Induction Labs is arguing that this requirement is the bottleneck.