Creator perspective · Money
Zero-order training doesn't scale: gradient-estimate error grows linearly with model size
A key limitation of the SPSA-style solver: the error in the gradient estimate scales worse — linearly less accurate — as the model size grows. Because of this, you can't train big models with this zero-order optimizer: for a 10-billion-parameter model, the number of perturbations required would be extremely large and infeasible — certainly more flops than a forward pass plus a backward pass — so it gets really compute inefficient.
Y Combinator · What If We Stopped Using GPUs? | YC Paper Club
Claim from Y Combinator host
host's analysis