Reply 20 of 20, by Gahhhrrrlic
It's because everything that enters the FPU has to be loaded from memory (except copying FPU registers or using the constant instructions like FLDZ) so you have the obligatory 16 cycle minimum which depends on bus speed and then however long it takes to retrieve the data from wherever it happens to be. If you use stack in your code the value will likely end up in cache. If you do a direct memory load it may or may not be cached. If you have L1 cache at say 28ns instead of 40ns L2, then any hot code will avoid bus contention, wait states, dram flushes, etc. So it's not so much that the cpu improvement helps the fpu. It's that you get a better cache hit rate and reduced latency loading and storing to the fpu.
I agree this is probably difficult to test. You'd have to run a kernel for example that does nothing but constant loads (0, 1, pi, e, etc) and math on them, then do memory loads and math via stack, then via direct memory load to several cold locations in a circuit perhaps. That would cover your fpu internals, your L1/L2 and your RAM so you could discriminate the performance gap.